← all papers · overview

Ternarylm: Memory-efficient Language Modeling Via Native 1.5-bit Quantization With Adaptive Layer-wise Scaling

Abstract

Large language models (LLMs) achieve remarkable performance but demand substantial computational resources, limiting deployment on edge devices and resource-constrained environments. We present TernaryLM, a 132M-parameter transformer trained natively with ternary quantization \{-1, 0, +1\} (log2(3) ~ 1.58-bit effective precision), achieving significant memory reduction without sacrificing language

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).