← all papers · overview

Batquant: Outlier-resilient MXFP4 Quantization Via Learnable Block-wise Optimization

Abstract

Microscaling floating-point (MXFP) formats have emerged as a promising standard for deploying Multi-modal Large Language Models (MLLMs) and Large Language Models (LLMs) on modern accelerator architectures. However, existing Post-Training Quantization (PTQ) methods, particularly rotation-based techniques designed for integer formats, suffer from severe performance collapse when applied to MXFP4. Re

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).