BrowseComp
Emerging6papers using it
555HF downloads
0HF likes
2025first seen
BrowseComp is a benchmark used to evaluate the performance of models in managing context during multi-round interactions.
π€ Hugging Faceβ apache-2.0
Papers using BrowseComp (6)
- Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement LearningSlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward GatingOpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty TrajectoriesStep 3.5 Flash: Open Frontier-Level Intelligence with 11B Active ParametersA$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid ReasoningWebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning