← all papers · overview

DARG: Dynamic Evaluation Of Large Language Models Via Adaptive Reasoning Graph

Abstract

The current paradigm of evaluating Large Language Models (LLMs) through static benchmarks comes with significant limitations, such as vulnerability to data contamination and a lack of adaptability to the evolving capabilities of LLMs. Therefore, evaluation methods that can adapt and generate evaluation data with controlled complexity are urgently needed. In this work, we introduce Dynamic Evaluati

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).