← all papers · overview

Dispersion Loss Counteracts Embedding Condensation And Improves Generalization In Small Language Models

Abstract

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in the smaller models. We observe a geometric phenomenon w

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).