Abstract
We provide a reproducible, end-To-end demonstration of vector search with OpenAI embeddings using Lucene on the popular MS MARCO passage ranking test collection. The main goal of our work is to challenge the prevailing narrative that a dedicated vector store is necessary to take advantage of recent advances in deep neural networks as applied to search. Quite the contrary, we show that hierarchical navigable small-world network (HNSW) indexes in Lucene are adequate to provide vector search capabilities in a standard bi-encoder architecture. This suggests that, from a simple cost-benefit analysis, there does not appear to be a compelling reason to introduce a dedicated vector store into a modern "AI stack"for search, since such applications have already received substantial investments in existing, widely deployed infrastructure.
Author supplied keywords
Cite
CITATION STYLE
Xian, J., Teofili, T., Pradeep, R., & Lin, J. (2024). Vector Search with OpenAI Embeddings: Lucene Is All You Need. In WSDM 2024 - Proceedings of the 17th ACM International Conference on Web Search and Data Mining (pp. 1090–1093). Association for Computing Machinery, Inc. https://doi.org/10.1145/3616855.3635691
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.