HLE
Emerging9papers using it
2025first seen
[!NOTE] IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing the dataset. Humanity's Last Exam π Website | π Paper | GitHub Center for AI Safety & Scale AI Humanity's Last Exam (HLE) is a multi-modal benchmark at the frontier of human knowledge, desi
Papers using HLE (9)
- Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B AgentScaling Agents via Continual Pre-trainingSAM: State-Adaptive Memory for Long-Horizon Reasoning AgentSEAGym: An Evaluation Environment for Self-Evolving LLM AgentsSCOPE: Prompt Evolution For Enhancing Agent EffectivenessEvoMAS: Learning Execution-Time Workflows for Multi-Agent SystemsMiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research TasksMirothinker: Pushing The Performance Boundaries Of Open-source Research Agents Via Model, Context, And Interactive ScalingFlowSearch: Advancing deep research with dynamic structured knowledge flow