← all papers · overview

Mobie: Efficient Inference Of Mixture Of Binary Experts Under Post-training Quantization

Abstract

Mixture-of-Experts (MoE) based large language models (LLMs) offer strong performance but suffer from high memory and computation costs. Weight binarization provides extreme efficiency, yet existing binary methods designed for dense LLMs struggle with MoE-specific issues, including cross-expert redundancy, task-agnostic importance estimatio

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).