Billion-scale Similarity Search Using A Hybrid Indexing Approach With Advanced Filtering
2025 Β· Simeon Emanuilov, Aleksandar Dimov
Abstract
This paper presents a novel approach for similarity search with complex filtering capabilities on billion-scale datasets, optimized for CPU inference. Our method extends the classical IVF-Flat index structure to integrate multi-dimensional filters. The proposed algorithm combines dense embeddings with discrete filtering attributes, enabling fast retrieval in high-dimensional spaces. Designed specifically for CPU-based systems, our disk-based approach offers a cost-effective solution for large-scale similarity search. We demonstrate the effectiveness of our method through a case study, showcasing its potential for various practical uses.
Authors
(none)
Tags
Stats
Related papers
- Adaptive Prefiltering For High-dimensional Similarity Search: A Frequency-aware Approach (2025)0.00
- Billion-scale Similarity Search With Gpus (2017)24.96
- Hybrid Inverted Index Is A Robust Accelerator For Dense Retrieval (2022)7.07
- Revisiting The Inverted Indices For Billion-scale Approximate Nearest Neighbors (2018)13.60
- Starling: An I/o-efficient Disk-resident Graph Index Framework For High-dimensional Vector Similarity Search On Data Segment (2024)12.61
- All-in-one Graph-based Indexing For Hybrid Search On Gpus (2025)0.00
- Hd-index: Pushing The Scalability-accuracy Boundary For Approximate Knn Search In High-dimensional Spaces (2018)14.02
- Embassi: Embedding Assignment Costs For Similarity Search In Large Graph Databases (2021)2.26