Self-Supervised CLIP-Based Image Recognition and Analysis for Electronic Data Forensics

1Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

To enhance intelligent identification of image authenticity and tampering in electronic data forensics, this paper proposes a self-supervised CLIP-based image recognition and analysis framework. Addressing current forensic limitations such as reliance on manual expertise and insufficient multimodal semantic understanding, the framework adopts the CLIP model as the foundation for image–text semantic alignment. By integrating a cross-attention mechanism, it strengthens fine-grained semantic association modeling between images and text. In addition, multiple self-supervised learning tasks, including image contrastive learning and pseudo-label prediction, are designed to guide the model in learning potential tampering features and semantic deviations under unsupervised conditions. Experiments conducted on datasets containing both authentic and manipulated images demonstrate that the proposed method achieves superior accuracy, robustness, and interpretability in both authenticity assessment and suspicious region localization. Compared with traditional supervised approaches and mainstream deep learning methods, the framework shows notable advantages. This work not only presents broad application potential in electronic data forensics but also provides a practical foundation for advancing multimodal learning within judicial technology.

Cite

CITATION STYLE

APA

Wang, Y. (2026). Self-Supervised CLIP-Based Image Recognition and Analysis for Electronic Data Forensics. IEEE Access, 14, 3062–3077. https://doi.org/10.1109/ACCESS.2025.3648849

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free