← all papers · overview

Diagonal-tiled Mixed-precision Attention For Efficient Low-bit MXFP Inference

Abstract

Transformer-based large language models (LLMs) have demonstrated remarkable performance across a wide range of real-world tasks, but their inference cost remains prohibitively high due to the quadratic complexity of attention and the memory bandwidth limitations of high-precision operations. In this work, we present a low-bit mixed-precision attention kernel using the microscaling floating-point (

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).