← all papers · overview

Mlora: Fine-tuning Lora Adapters Via Highly-efficient Pipeline Parallelism In Multiple Gpus

Abstract

Transformer-based, pre-trained large language models (LLMs) have demonstrated outstanding performance across diverse domains, particularly in the emerging \{\em pretrain-then-finetune\} paradigm. Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning method, is commonly used to adapt a base LLM to multiple downstream tasks. Further, LLM platforms enable developers to fine-tune multiple mode

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).