Evaluating information-retrieval models and machine-learning classifiers for measuring the social perception towards infectious diseases

11Citations
Citations of this article
40Readers
Mendeley users who have this article in their library.

Abstract

Recent outbreaks of infectious diseases remind us the importance of early-detection systems improvement. Infodemiology is a novel research field that analyzes online information regarding public health that aims to complement traditional surveillance methods. However, the large volume of information requires the development of algorithms that handle natural language efficiently. In the bibliography, it is possible to find different techniques to carry out these infodemiology studies. However, as far as our knowledge, there are no comprehensive studies that compare the accuracy of these techniques. Consequently, we conducted an infodemiology-based study to extract positive or negative utterances related to infectious diseases so that future syndromic surveillance systems can be improved. The contribution of this paper is two-fold. On the one hand, we use Twitter to compile and label a balanced corpus of infectious diseases with 6164 utterances written in Spanish and collected from Central America. On the other hand, we compare two statistical-models: word-grams and char-grams. The experimentation involved the analysis of different gram sizes, different partitions of the corpus, and two machine-learning classifiers: Random-Forest and Sequential Minimal Optimization. The results reach a 90.80% of accuracy applying the char-grams model with five-char-gram sequences. As a final contribution, the compiled corpus is released.

Cite

CITATION STYLE

APA

Apolinardo-Arzube, O., García-Díaz, J. A., Medina-Moreira, J., Luna-Aveiga, H., & Valencia-García, R. (2019). Evaluating information-retrieval models and machine-learning classifiers for measuring the social perception towards infectious diseases. Applied Sciences (Switzerland), 9(14). https://doi.org/10.3390/app9142858

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free