← all papers · overview

Bild: Bi-directional Logits Difference Loss For Large Language Model Distillation

Abstract

In recent years, large language models (LLMs) have shown exceptional capabilities across various natural language processing (NLP) tasks. However, such impressive performance often comes with the trade-off of an increased parameter size, posing significant challenges for widespread deployment. Knowledge distillation (KD) provides a solution by transferring knowledge from a large teacher model to a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).