Humanity's Last Exam
Emerging8papers using it
2025first seen
'Humanity's Last Exam' is a benchmark used to evaluate the performance of large language models in long-horizon, deep information-seeking research tasks.
Papers using Humanity's Last Exam (7)
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data SynthesisEvoMaster: A Foundational Evolving Agent Framework for Agentic Science at ScaleTongyi DeepResearch Technical ReportReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence ControlYunque DeepResearch Technical ReportMonoScale: Scaling Multi-Agent System with Monotonic ImprovementToolorchestra: Elevating Intelligence Via Efficient Model And Tool Orchestration