Combining Lexical, Host, and Content-based features for Phishing Websites detection using Machine Learning Models

  • Hamadouche S
  • Boudraa O
  • Gasmi M
N/ACitations
Citations of this article
61Readers
Mendeley users who have this article in their library.

Abstract

In cybersecurity field, identifying and dealing with threats from malicious websites (phishing, spam, and drive-by downloads, for example) is a major concern for the community. Consequently, the need for effective detection methods has become a necessity. Recent advances in Machine Learning (ML) have renewed interest in its application to a variety of cybersecurity challenges. When it comes to detecting phishing URLs, machine learning relies on specific attributes, such as lexical, host, and content based features. The main objective of our work is to propose, implement and evaluate a solution for identifying phishing URLs based on a combination of these feature sets. This paper focuses on using a new balanced dataset, extracting useful features from it, and selecting the optimal features using different feature selection techniques to build and conduct acomparative performance evaluation of four ML models (SVM, Decision Tree, Random Forest, and XGBoost). Results showed that the XGBoost model outperformed the others models, with an accuracy of 95.70% and a false negatives rate of 1.94%.

Cite

CITATION STYLE

APA

Hamadouche, S., Boudraa, O., & Gasmi, M. (2024). Combining Lexical, Host, and Content-based features for Phishing Websites detection using Machine Learning Models. ICST Transactions on Scalable Information Systems. https://doi.org/10.4108/eetsis.4421

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free