← all papers · overview

Agibench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark For Large Language Models

Abstract

Large language models (LLMs) like ChatGPT have revealed amazing intelligence. How to evaluate the question-solving abilities of LLMs and their degrees of intelligence is a hot-spot but challenging issue. First, the question-solving abilities are interlaced with different ability branches like understanding and massive knowledge categories like mathematics. Second, the inputs of questions are multi

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).