BFCLv-3
Emerging14papers using it
2025first seen
BFCLv-3 is a benchmark dataset used to evaluate the effectiveness of large language model agents in generating and refining tool calls through simulated feedback.
Papers using BFCLv-3 (13)
- D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool UsePushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and ActivationMAVEN: Improving Generalization in Agentic Tool CallingTopoCurate:Modeling Interaction Topology for Tool-Use Agent TrainingToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling DialoguesHINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon AgentsEnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RLControllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement LearningMagicAgent: Towards Generalized Agent PlanningGecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool CallsOn Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN DatasetSmall Language Models For Agentic Systems: A Survey Of Architectures, Capabilities, And Deployment Trade OffsToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning