Qwen-3 8B
Emerging3papers using it
2026first seen
The 'Qwen-3-8B' dataset/benchmark contains diverse, synthesized agentic tasks derived from real-world tool use, and it is used to evaluate the generalization capabilities of post-training tool-using large language models (LLMs) under varying task and toolset conditions.