← all papers · overview

Patch The Distribution Mismatch: RL Rewriting Agent For Stable Off-policy SFT

Abstract

Large language models (LLMs) have made rapid progress, yet adapting them to downstream scenarios still commonly relies on supervised fine-tuning (SFT). When downstream data exhibit a substantial distribution shift from the model's prior training distribution, SFT can induce catastrophic forgetting. To narrow this gap, data rewriting has been proposed as a data-centric approach that rewrites downst

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).