Vector Search with OpenAI Embeddings: Lucene Is All You Need

35Citations
Citations of this article
63Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We provide a reproducible, end-To-end demonstration of vector search with OpenAI embeddings using Lucene on the popular MS MARCO passage ranking test collection. The main goal of our work is to challenge the prevailing narrative that a dedicated vector store is necessary to take advantage of recent advances in deep neural networks as applied to search. Quite the contrary, we show that hierarchical navigable small-world network (HNSW) indexes in Lucene are adequate to provide vector search capabilities in a standard bi-encoder architecture. This suggests that, from a simple cost-benefit analysis, there does not appear to be a compelling reason to introduce a dedicated vector store into a modern "AI stack"for search, since such applications have already received substantial investments in existing, widely deployed infrastructure.

Author supplied keywords

Cite

CITATION STYLE

APA

Xian, J., Teofili, T., Pradeep, R., & Lin, J. (2024). Vector Search with OpenAI Embeddings: Lucene Is All You Need. In WSDM 2024 - Proceedings of the 17th ACM International Conference on Web Search and Data Mining (pp. 1090–1093). Association for Computing Machinery, Inc. https://doi.org/10.1145/3616855.3635691

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free