← all papers · overview

Lora-switch: Boosting The Efficiency Of Dynamic LLM Adapters Via System-algorithm Co-design

Abstract

Recent literature has found that an effective method to customize or further improve large language models (LLMs) is to add dynamic adapters, such as low-rank adapters (LoRA) with Mixture-of-Experts (MoE) structures. Though such dynamic adapters incur modest computational complexity, they surprisingly lead to huge inference latency overhead, slowing down the decoding speed by 2.5+ times. In this p

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).