Learning Sparse Lexical Representations Over Specified Vocabularies for Retrieval

6Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

A recent line of work in first-stage Neural Information Retrieval has focused on learning sparse lexical representations instead of dense embeddings. One such work is SPLADE, which has been shown to lead to state-of-the-art results in both the in-domain and zero-shot settings, can leverage inverted indices for efficient retrieval, and offers enhanced interpretability. However, existing SPLADE models are fundamentally limited to learning a sparse representation based on the native BERT WordPiece vocabulary. In this work, we extend SPLADE to support learning sparse representations over arbitrary sets of tokens to improve flexibility and aid integration with existing retrieval systems. As an illustrative example, we focus on learning a sparse representation over a large (300k) set of unigrams. We add an unsupervised pretraining task on C4 to learn internal representations for new tokens. Our experiments show that our Expanded-SPLADE model maintains the performance of WordPiece-SPLADE on both in-domain and zero-shot retrieval while allowing for custom output vocabularies.

Cite

CITATION STYLE

APA

Dudek, J. M., Kong, W., Li, C., Zhang, M., & Bendersky, M. (2023). Learning Sparse Lexical Representations Over Specified Vocabularies for Retrieval. In International Conference on Information and Knowledge Management, Proceedings (pp. 3865–3869). Association for Computing Machinery. https://doi.org/10.1145/3583780.3615207

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free