← all papers · overview

Revisiting On-policy Distillation: Empirical Failure Modes And Simple Fixes

Abstract

On-policy distillation (OPD) is appealing for large language model (LLM) post-training because it evaluates teacher feedback on student-generated rollouts rather than fixed teacher traces. In long-horizon settings, however, the common sampled-token variant is fragile: it reduces distribution matching to a one-token signal and becomes increasingly unreliable as rollouts drift away from prefixes the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).