← all papers · overview

Dual-space Knowledge Distillation For Large Language Models

Abstract

Knowledge distillation (KD) is known as a promising solution to compress large language models (LLMs) via transferring their knowledge to smaller models. During this process, white-box KD methods usually minimize the distance between the output distributions of the two models so that more knowledge can be transferred. However, in the current white-box KD framework, the output distributions are fro

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).