← all papers · overview

A Deep Dive Into The Trade-offs Of Parameter-efficient Preference Alignment Techniques

Abstract

Large language models are first pre-trained on trillions of tokens and then instruction-tuned or aligned to specific preferences. While pre-training remains out of reach for most researchers due to the compute required, fine-tuning has become affordable thanks to parameter-efficient methods such as LoRA and QLoRA. Alignment is known to be sensitive to the many factors involved, including the quant

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).