Abstract
We compare the multinomial i-vector framework from the speech community with LDA, SAGE, and LSA as feature learners for topic ID on multinomial speech and text data. We also compare the learned representations in their ability to discover topics, quantified by distributional similarity to gold-standard topics and by human interpretability. We find that topic ID and topic discovery are competing objectives. We argue that LSA and i-vectors should be more widely considered by the text processing community as pre-processing steps for downstream tasks, and also speculate about speech processing tasks that could benefit from more interpretable representations like SAGE.
Cite
CITATION STYLE
May, C., Ferraro, F., McCree, A., Wintrode, J., Garcia-Romero, D., & Van Durme, B. (2015). Topic identification and discovery on text and speech. In Conference Proceedings - EMNLP 2015: Conference on Empirical Methods in Natural Language Processing (pp. 2377–2387). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/d15-1285
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.