Automatic Generation of Topic Labels

Areej Alokaili; Nikolaos Aletras; Mark Stevenson

Conference ProceedingsOPEN ACCESS

Automatic Generation of Topic Labels

SIGIR 2020 - Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (2020) 1965-1968

DOI: 10.1145/3397271.3401185

16Citations

47Readers

Get full text

Abstract

Topic modelling is a popular unsupervised method for identifying the underlying themes in document collections that has many applications in information retrieval. A topic is usually represented by a list of terms ranked by their probability but, since these can be difficult to interpret, various approaches have been developed to assign descriptive labels to topics. Previous work on the automatic assignment of labels to topics has relied on a two-stage approach: (1) candidate labels are retrieved from a large pool (e.g. Wikipedia article titles); and then (2) re-ranked based on their semantic similarity to the topic terms. However, these extractive approaches can only assign candidate labels from a restricted set that may not include any suitable ones. This paper proposes using a sequence-to-sequence neural-based approach to generate labels that does not suffer from this limitation. The model is trained over a new large synthetic dataset created using distant supervision. The method is evaluated by comparing the labels it generates to ones rated by humans.

Author supplied keywords

Cite

CITATION STYLE

APA

Alokaili, A., Aletras, N., & Stevenson, M. (2020). Automatic Generation of Topic Labels. In SIGIR 2020 - Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 1965–1968). Association for Computing Machinery, Inc. https://doi.org/10.1145/3397271.3401185

Automatic Generation of Topic Labels

Abstract

Author supplied keywords

Cite

Register to see more suggestions