Advance Fake Video Detection via Vision Transformers

7Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Recent advancements in AI-based multimedia generation have enabled the creation of hyper-realistic images and videos, raising concerns about their potential use in spreading misinformation. The widespread accessibility of generative techniques, which allow for the production of fake multimedia from prompts or existing media, along with their continuous refinement, underscores the urgent need for highly accurate and generalizable AI-generated media detection methods, underlined also by new regulations like the European Digital AI Act. In this paper, we draw inspiration from Vision Transformer (ViT)-based fake image detection and extend this idea to video. We propose an original framework that effectively integrates ViT embeddings over time to enhance detection performance. Our method shows promising accuracy, generalization, and few-shot learning capabilities across a new, large and diverse dataset of videos generated using five open source generative techniques from the state-of-the-art, as well as a separate dataset containing videos produced by proprietary generative methods.

Cite

CITATION STYLE

APA

Battocchio, J., Dell’anna, S., Montibeller, A., & Boato, G. (2025). Advance Fake Video Detection via Vision Transformers. In IHandMMSec 2025 - Proceedings of the 2025 ACM Workshop on Information Hiding and Multimedia Security (pp. 1–11). Association for Computing Machinery, Inc. https://doi.org/10.1145/3733102.3733129

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free