Abstract
Many strategies have been put forward to assess the credibility of online social media content, however, none of them focuses on the issue of accuracy paradox which mostly occurs in highly skewed datasets, a case that usually arises in real-life situations. The purpose of this paper is to explore the use of various machine learning models including Gaussian Naïve Bayes, Latent Dirichlet Allocation (LDA), Linear Regression, Logistic Regression, and Support Vector Machine (SVM) for identifying the credibility of tweets. This includes proposing a new algorithm where the generative properties of Gaussian naïve Bayes are integrated with the discriminative properties of logistic regression and the author evaluates its performance in terms of accuracy and prediction power of determining tweet credibility. The Machine Learning Models used in this study, implemented on the Twitter datasets extracted from various real-world events are compared based on their accuracy and predictive power, in determining the credibility of tweets, to identify various accuracy paradox cases. The proposed algorithm is then used for the credibility inference of tweets and the reduction in the number of accuracy paradox cases is monitored. An extensive experimental study is performed to evaluate the performance of the proposed model on Twitter datasets with varied degrees of skewness. Our proposed model achieved accuracy and predictive power of 97% and 94% for a balanced dataset and 99% and 93% for an imbalanced dataset with 99% skewness.
Author supplied keywords
Cite
CITATION STYLE
Basharat, S., Afzal, S., Bamhdi, A. M., Khurshid, S., & Chachoo, M. (2023). Predicting and Mitigating the Effect of Skewness on Credibility Assessment of Social Media Content Using Machine Learning: A Twitter Case Study. International Journal of Computer Theory and Engineering, 15(3), 101–110. https://doi.org/10.7763/IJCTE.2023.V15.1338
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.