← all papers · overview

Milora: Efficient Mixture Of Low-rank Adaptation For Large Language Models Fine-tuning

Abstract

Low-rank adaptation (LoRA) and its mixture-of-experts (MOE) variants are highly effective parameter-efficient fine-tuning (PEFT) methods. However, they introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules in the Transformer layer. To address this issue, we propose Mixture of Low-Rank Adaptation (MiLoRA), a novel and efficie

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).