BrowseComp-ZH
Emerging13papers using it
2025first seen
π§ BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by Brow
Papers using BrowseComp-ZH (12)
- SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep ResearchScaling Agents via Continual Pre-trainingSAM: State-Adaptive Memory for Long-Horizon Reasoning AgentBeyond Turn Limits: Training Deep Search Agents with Dynamic Context WindowMach-Mind-4-Flash Technical ReportTongyi DeepResearch Technical ReportInfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information SeekingMiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research TasksWebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web ReasoningYunque DeepResearch Technical ReportMARINE: Theoretical Optimization and Design for Multi-Agent Recursive IN-context EnhancementMirothinker: Pushing The Performance Boundaries Of Open-source Research Agents Via Model, Context, And Interactive Scaling