← all papers · overview

Minitron-ssm: Efficient Hybrid Language Model Compression Through Group-aware SSM Pruning

Abstract

Hybrid LLM architectures that combine Attention and State Space Models (SSMs) achieve state-of-the-art accuracy and runtime performance. Recent work has demonstrated that applying compression and distillation to Attention-only models yields smaller, more accurate models at a fraction of the training cost. In this work, we explore the effectiveness of compressing Hybrid architectures. We introduce

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).