← all papers · overview

Enhancing Knowledge Distillation Of Large Language Models Through Efficient Multi-modal Distribution Alignment

Abstract

Knowledge distillation (KD) is an effective model compression method that can transfer the internal capabilities of large language models (LLMs) to smaller ones. However, the multi-modal probability distribution predicted by teacher LLMs causes difficulties for student models to learn. In this paper, we first demonstrate the importance of multi-modal distribution alignment with experiments and the

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).