← all papers · overview

Universal Cross-tokenizer Distillation Via Approximate Likelihood Matching

Abstract

Distillation has shown remarkable success in transferring knowledge from a Large Language Model (LLM) teacher to a student LLM. However, current distillation methods require similar tokenizers between the teacher and the student, restricting their applicability to only a small subset of teacher-student pairs. In this work, we develop a principled cross-tokenizer distillation method to solve this c

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).