Search-adaptor: Embedding Customization For Information Retrieval
2023 Β· Jinsung Yoon, Sercan O Arik, Yanfei Chen, et al.
Abstract
Embeddings extracted by pre-trained Large Language Models (LLMs) have significant potential to improve information retrieval and search. Beyond the zero-shot setup in which they are being conventionally used, being able to take advantage of the information from the relevant query-corpus paired data can further boost the LLM capabilities. In this paper, we propose a novel method, Search-Adaptor, for customizing LLMs for information retrieval in an efficient and robust way. Search-Adaptor modifies the embeddings generated by pre-trained LLMs, and can be integrated with any LLM, including those only available via prediction APIs. On multiple English, multilingual, and multimodal retrieval datasets, we show consistent and significant performance benefits for Search-Adaptor -- e.g., more than 5% improvements for Google Embedding APIs in nDCG@10 averaged over 14 BEIR datasets.
Authors
(none)
Tags
Stats
Related papers
- Llm-augmented Retrieval: Enhancing Retrieval Models Through Language Models And Doc-level Embedding (2024)0.00
- Matryoshka-adaptor: Unsupervised And Supervised Tuning For Smaller Embedding Dimensions (2024)2.26
- Federated Learning With Ad-hoc Adapter Insertions: The Case Of Soft-embeddings For Training Classifier-as-retriever (2025)0.00
- LMAR: Language Model Augmented Retriever For Domain-specific Knowledge Indexing (2025)1.57
- Large Reasoning Embedding Models: Towards Next-generation Dense Retrieval Paradigm (2025)0.00
- Modernizing Facebook Scoped Search: Keyword And Embedding Hybrid Retrieval With LLM Evaluation (2025)0.00
- Mm-embed: Universal Multimodal Retrieval With Multimodal Llms (2024)0.00
- Align Then Train: Efficient Retrieval Adapter Learning (2026)0.00