COVID-19 Dataset Clustering based on K-Means and EM Algorithms

9Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

In this paper, a COVID-19 dataset is analyzed using a combination of K-Means and Expectation-Maximization (EM) algorithms to cluster the data. The purpose of this method is to gain insight into and interpret the various components of the data. The study focuses on tracking the evolution of confirmed, death, and recovered cases from March to October 2020, using a two-dimensional dataset approach. K-Means is used to group the data into three categories: “Confirmed-Recovered”, “Confirmed-Death”, and “Recovered-Death”, and each category is modeled using a bivariate Gaussian density. The optimal value for k, which represents the number of groups, is determined using the Elbow method. The results indicate that the clusters generated by K-Means provide limited information, whereas the EM algorithm reveals the correlation between “Confirmed-Recovered”, “Confirmed-Death”, and “Recovered-Death”. The advantages of using the EM algorithm include stability in computation and improved clustering through the Gaussian Mixture Model (GMM).

Cite

CITATION STYLE

APA

Boutazart, Y., Satori, H., Affane, M. A. R., Hamidi, M., & Satori, K. (2023). COVID-19 Dataset Clustering based on K-Means and EM Algorithms. International Journal of Advanced Computer Science and Applications, 14(3), 924–934. https://doi.org/10.14569/IJACSA.2023.01403105

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free