Beyond Hard Negatives: The Importance Of Score Distribution In Knowledge Distillation For Dense Retrieval

Abstract

arXiv:2604.04734v2 Announce Type: replace Abstract: Transferring knowledge from a cross-encoder teacher via Knowledge Distillation (KD) has become a standard paradigm for training retrieval models. While existing studies have largely focused on mining hard negatives to improve discrimination, the systematic composition of training data and the resulting teacher score distribution have received relatively less attention. In this work, we highlight that focusing solely on hard negatives prevents the student from learning the comprehensive preference structure of the teacher, potentially hampering generalization. To effectively emulate the teacher score distribution, we propose a Stratified Sampling strategy that uniformly covers the entire score spectrum. Experiments on in-domain and out-of-domain benchmarks confirm that Stratified Sampling, which preserves the variance and entropy of teacher scores, serves as a robust baseline, significantly outperforming top-K and random sampling in d

Beyond Hard Negatives: The Importance Of Score Distribution In Knowledge Distillation For Dense Retrieval

Abstract

Authors

Tags

Stats

Related papers