Integrated Named Entity Recognition and Identical-Entity Detection for Extracting Unique Information Sources in News Articles

  • Ansyah A
  • Oranova Siahaan D
  • Izqi Paradisiaca B
N/ACitations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

Native advertising is often difficult to detect because it resembles regular news articles. One indicator is the absence of diverse information sources or the reliance on a single perspective. Therefore, it is necessary to employ an extraction technique capable of consolidating various forms of identical entity mentions. This study integrates an NER model based on XLNet+BiLSTM+CRF with identical entity classification using Levenshtein distance features and static and contextual vector representations. The results show an F1-score of 93.71% at the entity level and 92.84% for identical entity identification, along with a list of unique citation sources. These findings demonstrate that this unique list can be an additional feature in detecting native advertising, which often relies on a single source. With an average unique entity coverage of 97.40%, the proposed architecture can extract unique entities within news articles

Cite

CITATION STYLE

APA

Ansyah, A. S. S., Oranova Siahaan, D., & Izqi Paradisiaca, B. R. (2025). Integrated Named Entity Recognition and Identical-Entity Detection for Extracting Unique Information Sources in News Articles. Digital Zone: Jurnal Teknologi Informasi Dan Komunikasi, 16(2), 72–83. https://doi.org/10.31849/digitalzone.v16i2.27687

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free