Finding common features in multilingual fake news: a quantitative clustering approach

Wei Yuan; Haitao Liu

Journal Article

Finding common features in multilingual fake news: a quantitative clustering approach

Digital Scholarship in the Humanities (2024) 39(2) 790-804

DOI: 10.1093/llc/fqae016

3Citations

5Readers

Get full text

Abstract

Since the Internet is a breeding ground for unconfirmed fake news, its automatic detection and clustering studies have become crucial. Most current studies focus on English texts, and the common features of multilingual fake news are not sufficiently studied. Therefore, this article uses English, Russian, and Chinese as examples and focuses on identifying the common quantitative features of fake news in different languages at the word, sentence, readability, and sentiment levels. These features are then utilized in principal component analysis, K-means clustering, hierarchical clustering, and two-step clustering experiments, which achieved satisfactory results. The common features we proposed play a greater role in achieving automatic cross-lingual clustering than the features proposed in previous studies. Simultaneously, we discovered a trend toward linguistic simplification and economy in fake news. Furthermore, fake news is easier to understand and uses negative emotional expressions in ways that real news does not. Our research provides new reference features for fake news detection tasks and facilitates research into their linguistic characteristics.

Author supplied keywords

Cite

CITATION STYLE

APA

Yuan, W., & Liu, H. (2024). Finding common features in multilingual fake news: a quantitative clustering approach. Digital Scholarship in the Humanities, 39(2), 790–804. https://doi.org/10.1093/llc/fqae016

Finding common features in multilingual fake news: a quantitative clustering approach

Abstract

Author supplied keywords

Cite

Register to see more suggestions