← all papers · overview

Task-specific Knowledge Distillation Via Intermediate Probes

Abstract

Knowledge distillation from large language models (LLMs) assumes that the teacher's output distribution is a high-quality training signal. On reasoning tasks, this assumption is frequently violated. A model's intermediate representations may encode the correct answer, yet this information is lost or distorted through the vocabulary projection, where prompt formatting and answer-token choices creat

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).