← all papers · overview

M2R2: Mixture Of Multi-rate Residuals For Efficient Transformer Inference

Abstract

Residual transformations enhance the representational depth and expressive power of large language models (LLMs). However, applying static residual transformations across all tokens in auto-regressive generation leads to a suboptimal trade-off between inference efficiency and generation fidelity. Existing methods, including Early Exiting, Skip Decoding, and Mixture-of-Depth address this by modulat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).