← all datasets

BrowseComp-ZH

Emerging
13papers using it
2025first seen

🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by Brow

Papers using BrowseComp-ZH (12)

BrowseComp-ZH dataset β€” papers, benchmarks & downloads Β· AI Agents