Toward Robust Malware Detection: A Survey of Datasets, Techniques, and Practical Challenges

  • Huynh T
  • Huynh D
  • Trinh V
N/ACitations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

The increasing sophistication of malware has diminished the effectiveness of traditional signature-based detection. While Machine Learning (ML), Deep Learning (DL), and Large Language Models (LLMs) have improved malware classification, real-world systems continue to struggle with evasion attacks, temporal drift, and class imbalance. This study reviews the advancements in robust malware detection, focusing on benchmark datasets, detection methods, and operational constraints. Public datasets - EMBER2018, SOREL-20M, MalDICT, MOTIF, and EMBER2024 - are assessed for scale, label quality, and reproducibility. This paper contributes: (i) a Robust Malware Evaluation Protocol (RMEP) for consistent benchmarking under low False-Positive Rates (FPR) (≤ 0.1%) with temporal splits, and (ii) a Dataset-Task-Robustness (DTR) matrix for systematic comparison, offering practical guidance for reproducible malware-detection research. Future efforts should focus on broader multi-platform benchmark coverage, explicit analysis of robustness–accuracy trade-offs, interpretable language-assisted detection pipelines, and privacy-preserving collaborative learning frameworks.

Cite

CITATION STYLE

APA

Huynh, T.-T., Huynh, D.-T., & Trinh, V.-Q. (2026). Toward Robust Malware Detection: A Survey of Datasets, Techniques, and Practical Challenges. Engineering, Technology & Applied Science Research, 16(3), 35064–35070. https://doi.org/10.48084/etasr.17500

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free