← all papers · overview

Fast And Accurate Probing Of In-training Llms' Downstream Performances

Abstract

The paradigm of scaling Large Language Models (LLMs) in both parameter size and test time has pushed the boundaries of AI capabilities, but at the cost of making the traditional generative evaluation paradigm prohibitively expensive, therefore making the latency of LLM's in-training downstream performance evaluation unbearable. However, simple metrics like training loss (perplexity) are not always

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).