← all papers · overview

Hybrid Gated Flow (HGF): Stabilizing 1.58-bit Llms Via Selective Low-rank Correction

Abstract

The deployment of Large Language Models (LLMs) on edge devices is fundamentally constrained by the "Memory Wall" -- a hardware limitation where memory bandwidth, not compute, becomes the bottleneck. Recent 1.58-bit quantization techniques (e.g., BitNet b1.58) dramatically reduce memory footprint but typically incur a perplexity degradation of 20-25% compared to FP16 baselines. In this work, we int

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).