MICE: Minimal Interaction Cross-encoders For Efficient Re-ranking
2026 Β· Mathias Vast, Victor Morand, Basile van Cooten, et al.
Abstract
Cross-encoders deliver state-of-the-art ranking effectiveness in information retrieval, but have a high inference cost. This prevents them from being used as first-stage rankers, but also incurs a cost when re-ranking documents. Prior work has addressed this bottleneck from two largely separate directions: accelerating cross-encoder inference by sparsifying the attention process or improving first-stage retrieval effectiveness using more complex models, e.g. late-interaction ones. In this work, we propose to bridge these two approaches, based on an in-depth understanding of the internal mechanisms of cross-encoders. Starting from cross-encoders, we show that it is possible to derive a new late-interaction-like architecture by carefully removing detrimental or unnecessary interactions. We name this architecture MICE (Minimal Interaction Cross-Encoders). We extensively evaluate MICE across both in-domain (ID) and out-of-domain (OOD) datasets. MICE decreases fourfold the inference latency
Authors
(none)
Tags
Stats
Related papers
- Efficient Document Ranking With Learnable Late Interactions (2024)0.00
- Efficient Neural Ranking Using Forward Indexes And Lightweight Encoders (2023)5.24
- CODER: An Efficient Framework For Improving Retrieval Through Contextual Document Embedding Reranking (2021)7.16
- Drowning In Documents: Consequences Of Scaling Reranker Inference (2024)0.00
- What Drives Cross-lingual Ranking? Retrieval Approaches With Multilingual Language Models (2025)0.00
- Mixlm: High-throughput And Effective LLM Ranking Via Text-embedding Mix-interaction (2025)0.00
- Shallow Cross-encoders For Low-latency Retrieval (2024)2.26
- Efficient Discriminative Joint Encoders For Large Scale Vision-language Reranking (2025)0.00