Exploring techniques to distinguish between real images and those generated using stable diffusion XL

N/ACitations
Citations of this article
4Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The recent development of text-to-image diffusion models has allowed us to quickly generate realistic images from textual prompts. Despite enabling innovation in particular domains, concerns have been raised over the prospect of malicious users posing synthetic images as genuine. To assess if it is possible to discern between real images and those generated using diffusion models, a novel convolutional neural network was built, trained and tested on a bespoke dataset formed of authentic images from the ImageNet dataset and corresponding synthetic images generated using Stable Diffusion XL: an open-source text-to-image diffusion model. With the public release of this dataset, it is currently the largest publicly accessible collection of images generated using Stable Diffusion XL, significantly contributing to future research in this area. The positive results from our experiment performing a binary classification of synthetic and real images demonstrate the effectiveness in detecting synthetic images, with up to 98.38% accuracy using a ResNet-18 baseline, and 97.24% with the proposed CNN.

Cite

CITATION STYLE

APA

Sanders, B., Morrison, D., & Harris-Birtill, D. (2026). Exploring techniques to distinguish between real images and those generated using stable diffusion XL. PLOS ONE, 21(1 January). https://doi.org/10.1371/journal.pone.0339917

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free