← all papers · overview

Adafuse: Accelerating Dynamic Adapter Inference Via Token-level Pre-gating And Fused Kernel Optimization

Abstract

The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Models (LLMs). However, this architectural enhancement comes at a steep cost: despite minimal increases in computational load, the inference latency often skyrockets, leading to decoding speeds slowing by over 2.5 times. Thr

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).