Enhancing Spam Email Detection with Machine Learning: A Comparative Study of Logistic Regression and Naive Bayes Using Apache Spark

  • Ye Z
N/ACitations
Citations of this article
13Readers
Mendeley users who have this article in their library.

Abstract

The spread of spam emails presents serious problems for both email security and user experience. This research aims to develop an effective spam email classification system utilizing machine learning techniques, specifically Logistic Regression and Naive Bayes, within the Apache Spark framework. The methodology encompasses a thorough preprocessing of the Enron email dataset. This process involves several critical steps: text cleaning to remove irrelevant information, tokenization to break down the text into individual words, removal of stop words to eliminate common but uninformative words, and text feature extraction using Term Frequency-Inverse Document Frequency (TF-IDF) to quantify the importance of terms within the dataset. The study is conducted on a subset of the Enron email dataset, comprising 11,029 emails, with 2,996 labeled as spam. Experimental results demonstrate that the Naive Bayes model outperforms Logistic Regression, achieving higher accuracy and F1 score. This finding underscores the robustness of Naive Bayes in spam email classification, highlighting its potential for enhancing email security by effectively filtering spam.

Cite

CITATION STYLE

APA

Ye, Z. (2024). Enhancing Spam Email Detection with Machine Learning: A Comparative Study of Logistic Regression and Naive Bayes Using Apache Spark. Transactions on Computer Science and Intelligent Systems Research, 7, 78–85. https://doi.org/10.62051/gt8zn492

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free