← all papers · overview

Model Utility Law: Evaluating Llms Beyond Performance Through Mechanism Interpretable Metric

Abstract

Large Language Models (LLMs) have become indispensable across academia, industry, and daily applications, yet current evaluation methods struggle to keep pace with their rapid development. One core challenge of evaluation in the large language model (LLM) era is the generalization issue: how to infer a model's near-unbounded abilities from inevitably bounded benchmarks. We address this challenge b

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).