Synthetic Malware Using Deep Variational Autoencoders and Generative Adversarial Networks

2Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

The effectiveness of detecting malicious files heavily relies on the quality of the training dataset, particularly its size and authenticity. However, the lack of high-quality training data remains one of the biggest challenges in achieving widespread adoption of malware detection by trained machine and deep learning models. In response to this challenge, researchers have made initial strides by employing generative techniques to create synthetic malware samples. This work utilizes deep variational autoencoders (VAE) and generative adversarial networks (GAN) to produce malware samples as opcode sequences. The generated malware opcodes are then distinguished from authentic opcode samples using machine and deep learning techniques as validation methods. The primary objective of this study was to compare synthetic malware generated using VAE and GAN technologies. The results showed that neither approach could create synthetic malware that could deceive machine learning classification. However, the WGAN-GP algorithm showed more promise by requiring a higher number of synthetic malware samples in the train set to effectively be detected, proving it a better approach in synthetic malware generation.

Author supplied keywords

Cite

CITATION STYLE

APA

Choi, A., Giang, A., Jumani, S., Luong, D., & Troia, F. D. (2024). Synthetic Malware Using Deep Variational Autoencoders and Generative Adversarial Networks. EAI Endorsed Transactions on Internet of Things, 10. https://doi.org/10.4108/eetiot.6566

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free