Automatic and human evaluation of local topic quality

9Citations
Citations of this article
128Readers
Mendeley users who have this article in their library.

Abstract

Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Recent models, which claim to improve token-level topic assignments, are only validated on global metrics. We elicit human judgments of token-level topic assignments: over a variety of topic model types and parameters, global metrics agree poorly with human assignments. Since human evaluation is expensive we propose automated metrics to evaluate topic models at a local level. Finally, we correlate our proposed metrics with human judgments: an evaluation based on the percent of topic switches correlates most strongly with human judgment of local topic quality. This new metric, which we call consistency, should be adopted alongside global metrics such as topic coherence.

Cite

CITATION STYLE

APA

Lund, J., Armstrong, P., Fearn, W., Cowley, S., Byun, C., Boyd-Graber, J., & Seppi, K. (2020). Automatic and human evaluation of local topic quality. In ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (pp. 788–796). Association for Computational Linguistics (ACL).

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free