Beyond Counting Words: Assessing Performance of Dictionaries, Supervised Machine Learning, and Embeddings in Topic and Frame Classif ication

20Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Topics and frames are at the heart of various theories in communication science and other social sciences, making their measurement of key interest to many scholars. The current study compares and contrasts two main deductive computational approaches to measure policy topics and frames: Dictionary (lexicon) based identif ication, and supervised machine learning. Additionally, we introduce domain-specif ic word embeddings to these classif ication tasks. Drawing on a manually coded dataset of Dutch news articles and parliamentary questions, our results indicate that super vised machine learning outperforms dictionar y-based classif ication for both tasks. Furthermore, results show that word embeddings may boost performance at relatively low cost by introducing relevant and domain-specif ic semantic information to the classif ication model.

Cite

CITATION STYLE

APA

Kroon, A. C., van der Meer, T., & Vliegenthart, R. (2022). Beyond Counting Words: Assessing Performance of Dictionaries, Supervised Machine Learning, and Embeddings in Topic and Frame Classif ication. Computational Communication Research, 4(2), 528–570. https://doi.org/10.5117/CCR2022.2.006.KROO

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free