Abstract
To enhance intelligent identification of image authenticity and tampering in electronic data forensics, this paper proposes a self-supervised CLIP-based image recognition and analysis framework. Addressing current forensic limitations such as reliance on manual expertise and insufficient multimodal semantic understanding, the framework adopts the CLIP model as the foundation for image–text semantic alignment. By integrating a cross-attention mechanism, it strengthens fine-grained semantic association modeling between images and text. In addition, multiple self-supervised learning tasks, including image contrastive learning and pseudo-label prediction, are designed to guide the model in learning potential tampering features and semantic deviations under unsupervised conditions. Experiments conducted on datasets containing both authentic and manipulated images demonstrate that the proposed method achieves superior accuracy, robustness, and interpretability in both authenticity assessment and suspicious region localization. Compared with traditional supervised approaches and mainstream deep learning methods, the framework shows notable advantages. This work not only presents broad application potential in electronic data forensics but also provides a practical foundation for advancing multimodal learning within judicial technology.
Author supplied keywords
Cite
CITATION STYLE
Wang, Y. (2026). Self-Supervised CLIP-Based Image Recognition and Analysis for Electronic Data Forensics. IEEE Access, 14, 3062–3077. https://doi.org/10.1109/ACCESS.2025.3648849
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.