← all papers · overview

Rethinking Perplexity: Revealing The Impact Of Input Length On Perplexity Evaluation In Llms

Abstract

Perplexity is a widely adopted metric for assessing the predictive quality of large language models (LLMs) and often serves as a reference metric for downstream evaluations. However, recent evidence shows that perplexity can be unreliable, especially when irrelevant long inputs are used, raising concerns for both benchmarking and system deployment. While prior efforts have employed selective input

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).