BrowseComp
Emerging21papers using it
2025first seen
BrowseComp is a benchmark that contains complex questions synthesized from live web traversal, used to evaluate the browsing competence of search agents by minimizing test-set contamination and parametric memorization.
Papers using BrowseComp (17)
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentFlash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel ExecutionSearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep ResearchLiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?EvoBrowseComp: Benchmarking Search Agents on Evolving KnowledgeSAM: State-Adaptive Memory for Long-Horizon Reasoning AgentEvoMaster: A Foundational Evolving Agent Framework for Agentic Science at ScaleSlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward GatingToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use CapabilityTongyi DeepResearch Technical ReportGeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning TracesW&D:Scaling Parallel Tool Calling for Efficient Deep Research AgentsSearch More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and GeneralizationWebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web ReasoningYunque DeepResearch Technical ReportMirothinker: Pushing The Performance Boundaries Of Open-source Research Agents Via Model, Context, And Interactive ScalingCOMPASS: Enhancing Agent Long-horizon Reasoning With Evolving Context