Long document ranking with query-directed sparse transformer

N/ACitations
Citations of this article
92Readers
Mendeley users who have this article in their library.

Abstract

The computing cost of transformer self-attention often necessitates breaking long documents to fit in pretrained models in document ranking tasks. In this paper, we design Query-Directed Sparse attention that induces IR-axiomatic structures in transformer self-attention. Our model, QDS-Transformer, enforces the principle properties desired in ranking: local contextualization, hierarchical representation, and query-oriented proximity matching, while it also enjoys efficiency from sparsity. Experiments on one fully supervised and three few-shot TREC document ranking benchmarks demonstrate the consistent and robust advantage of QDS-Transformer over previous approaches, as they either retrofit long documents into BERT or use sparse attention without emphasizing IR principles. We further quantify the computing complexity and demonstrates that our sparse attention with TVM implementation is twice more efficient that the fully-connected self-attention. All source codes, trained model, and predictions of this work are available at https://github.com/hallogameboy/ QDS-Transformer.

Cite

CITATION STYLE

APA

Jiang, J. Y., Xiong, C., Lee, C. J., & Wang, W. (2020). Long document ranking with query-directed sparse transformer. In Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020 (pp. 4594–4605). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.findings-emnlp.412

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free