Adversarial Attacks on Deepfake Detectors and Defence Mechanisms: A Cyber Security Model

9Citations
Citations of this article
21Readers
Mendeley users who have this article in their library.

Abstract

Deepfake detectors have grown increasingly important in ensuring digital content integrity, but they are prone to adversarial attacks aimed at manipulating the inputs with the intent of biasing the detection models. This work introduces an end-to-end methodology for enhancing robustness in deepfake detectors against adversarial threats. First, we discuss adversarial attack methods: Fast Gradient Sign Method and Projected Gradient Descent. Both generate adversarial examples by adding perturbations to the inputs in a way that misleads the detection system. Finally, in the presence of such adversarial attacks, we extend a multi-dimensional defense strategy entailing adversarial training, input preprocessing, defensive distillation, and randomized smoothing. Adversarial training strengthens the model by incorporating adversarial examples into training, whereas input preprocessing uses other techniques like denoising to filter out the perturbations. Defensive distillation strengthens model resilience by softening output predictions, while randomized smoothing averages predictions over noisy inputs and provides robustness to adversarial manipulation. This will fortify the defensive mechanism of deepfake detection systems and support the general adversarial defense realm in deep learning models. Both XceptionNet and MesoNet are widely recognized in the literature for their performance in image-based classification tasks, including deepfake detection. XceptionNet has a strong track record due to its depthwise separable convolutions, which offer computational efficiency without sacrificing accuracy. MesoNet, on the other hand, is specifically designed for detecting manipulated media, making it well-suited for this application. The proposed approach is compared against traditional deepfake detection methods and existing defense mechanisms to evaluate its effectiveness. Specifically: Baseline Deepfake Detection Models (XceptionNet, CNN-based architectures), Adversarial Attacks (FGSM, PGD) and Defense Mechanisms (adversarial training, gradient masking). The type of Dataset Used is DFDC (Deepfake Detection Challenge) Includes diverse face-swapping and manipulation techniques, making it suitable for evaluating deepfake detection under realistic conditions. The experimental results have shown that the proposed framework greatly enhances the robustness of deepfake detectors against adversarial attacks, thus ensuring more reliable performance in real-world applications. The proposed method achieves an accuracy of 94.1% with Adversarial Training and 92.8% with Randomized Smoothing.

Cite

CITATION STYLE

APA

Abed, B. N. A. din, Hussien, S. A. S., & Majeed, S. H. (2025). Adversarial Attacks on Deepfake Detectors and Defence Mechanisms: A Cyber Security Model. International Journal of Intelligent Engineering and Systems, 18(3), 730–745. https://doi.org/10.22266/ijies2025.0430.50

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free