Abstract
The rapid evolution of deep generative models has facilitated the creation of “Deepfakes”, enabling the synthesis of hyper-realistic facial manipulations that threaten the trustworthiness of digital media. While forensic countermeasures have been developed to identify these forgeries, deepfake detection in real-world scenarios is severely hampered by video compression artifacts, which often obscure the subtle pixel-level traces exploited by conventional Convolutional Neural Networks (CNNs). This study introduces a robust detection framework designed specifically to withstand the aggressive compression inherent to social media dissemination. We present a hybrid 3D architecture that integrates the local spatiotemporal feature extraction capabilities of a 3D-ResNet-50 backbone with the global context modeling of a temporal Video Vision Transformer. Unlike frame-based or joint spatiotemporal attention approaches, the proposed model performs fully video-level reasoning and utilizes a factorized self-attention mechanism to decouple spatial and temporal modeling, thereby preserving stable temporal cues under compression while minimizing computational costs. Experimental results on the compressed protocols of the FaceForensics++ dataset as well as Celeb-DF-v2 and DFDC datasets, including cross-dataset generalization evaluation, validate the efficacy of this design, demonstrating that our method achieves superior detection accuracy and generalization compared to existing baselines, particularly on low-quality inputs.
Author supplied keywords
Cite
CITATION STYLE
Saga, A., Lili, N. A., Khalid, F., Sani, N. F. M., Abdalla, H. E. M., Syahputra, Z., & Wijaya, R. F. (2025). Detecting Low-Quality Deepfake Videos Using 3D Residual Vision Transformer. International Journal of Advanced Computer Science and Applications, 16(12), 651–659. https://doi.org/10.14569/IJACSA.2025.0161260
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.