← all papers · overview

VISA: Value Injection Via Shielded Adaptation For Personalized LLM Alignment

Abstract

Aligning Large Language Models (LLMs) with nuanced human values remains a critical challenge, as existing methods like Reinforcement Learning from Human Feedback (RLHF) often handle only coarse-grained attributes. In practice, fine-tuning LLMs on task-specific datasets to optimize value alignment inevitably incurs an alignment tax: the model's pre-calibrated value system drifts significantly due t

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).