← all papers · overview

SNIP: An Adaptive Mixed Precision Framework For Subbyte Large Language Model Training

Abstract

Training large language models (LLMs) efficiently while preserving model quality poses significant challenges, particularly with subbyte precision supported by state-of-the-art GPUs. Current mixed-precision training approaches either apply uniform precision to all GEMM operations or rely on heuristic-based methods that fail to generalize during training, leading to suboptimal convergence and insta

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).