Dynamic Spatial-Temporal Inconsistency Learning for General Deepfake Detection in Visual Understanding

N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Generalizable deepfake detection is essential for trustworthy visual understanding in real-world computer vision applications. This paper presents a dynamic spatial-temporal inconsistency learning algorithm designed to achieve high generalization in deepfake video detection. Current video-based detection approaches tend to either isolate spatial artifacts or merely exploit coarse temporal inconsistencies when identifying deepfake videos, which impedes the acquisition of fine-grained spatial-temporal clues and consequently limits their generalization capability. To this end, we propose the dynamic spatial-temporal network (DST-Net), a deep architecture that systematically mines comprehensive inconsistency cues through three synergistic modules. The short-term temporal modality extraction (STME) module captures temporal dynamics from adjacent frames. The short-term spatial-temporal inconsistency extraction (SSTIE) module with pixel-wise supervision learns semantically meaningful inconsistency features resistant to perturbations. The dynamic-term spatial-temporal inconsistency extraction (DSTIE) module adaptively aggregates these features across timescales, building robust multi-scale representations. This design ensures that the learned representations capture intrinsic forgery patterns, enhancing generalization and robustness. Comprehensive evaluations conducted on five widely adopted benchmark datasets reveal that our method surpasses nine representative competitors, with superior robustness to common image perturbations. This work advances the application of deep learning algorithms to reliable visual understanding in multimedia forensics.

Cite

CITATION STYLE

APA

Li, J., Liao, G., Wang, Y., Liu, X., & Liu, B. (2026). Dynamic Spatial-Temporal Inconsistency Learning for General Deepfake Detection in Visual Understanding. Mathematics, 14(10). https://doi.org/10.3390/math14101612

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free