Humanity's Last Exam (HLE)
Emerging4papers using it
2025first seen
The Humanity's Last Exam (HLE) is a benchmark dataset used to evaluate the performance of models in solving deep and complex problems.
Papers using Humanity's Last Exam (HLE) (4)
- OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty TrajectoriesToolOrchestra: Elevating Intelligence via Efficient Model and Tool OrchestrationAgentic Entropy-Balanced Policy OptimizationSFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents