Wangiri Fraud Detection: A Comprehensive Approach to Unlabeled Telecom Data

1Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

Abstract

Wangiri fraud is a pervasive telecommunications scam that exploits missed calls to lure victims into dialing premium-rate numbers, resulting in significant financial losses for operators and consumers. This paper presents a comprehensive machine learning framework for detecting Wangiri fraud in highly imbalanced and unlabeled Call Detail Record (CDR) datasets. We introduce a novel unsupervised labeling approach using domain-driven heuristics, coupled with advanced feature engineering to capture temporal, geographic, and behavioral patterns indicative of fraud. To address severe class imbalance, we evaluate multiple sampling strategies like the Synthetic Minority Over-sampling Technique (SMOTE) and undersampling, and also compare the performance of Logistic Regression, Decision Trees, Random Forest, XGBoost, and Multi-Layer Perceptron (MLP). Our results demonstrate that ensemble methods, particularly Random Forest and XGBoost, achieve near-perfect accuracy (e.g., Receiver Operating Characteristic Area Under the Curve (ROC-AUC) (Formula presented.)) on balanced data while maintaining interpretability. The proposed pipeline offers a scalable and practical solution for real-time fraud detection, providing telecom operators with an effective tool to mitigate Wangiri fraud risks.

Cite

CITATION STYLE

APA

Balouchi, A., Abdollahi, M., Eskandarian, A., Karimi Pour Kerman, K., Majd, E., Azouji, N., & Baniasadi, A. (2026). Wangiri Fraud Detection: A Comprehensive Approach to Unlabeled Telecom Data. Future Internet, 18(1). https://doi.org/10.3390/fi18010015

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free