← all papers · overview

Researcharena: Benchmarking Large Language Models' Ability To Collect And Organize Information As Research Agents

Abstract

Large language models (LLMs) excel across many natural language processing tasks but face challenges in domain-specific, analytical tasks such as conducting research surveys. This study introduces ResearchArena, a benchmark designed to evaluate LLMs' capabilities in conducting academic surveys -- a foundational step in academic research. ResearchArena models the process in three stages: (1) inform

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).