← all papers · overview

Sparse Semantic Dimension As A Generalization Certificate For Llms

Abstract

Standard statistical learning theory predicts that Large Language Models (LLMs) should overfit because their parameter counts vastly exceed the number of training tokens. Yet, in practice, they generalize robustly. We propose that the effective capacity controlling generalization lies in the geometry of the model's internal representations: while the parameter space is high-dimensional, the activa

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).