A Human-factors Approach for Evaluating AI-generated Images

5Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

As generative artificial intelligence (AI) becomes more common in day-to-day life, AI-generated content (AIGC) needs to be accurate, relevant, and comprehensive. These characteristics typically are determined by subjective, human-based image quality assessment; however, there is limited research on the qualification of AI-generated image quality. Over 9,800 images were generated using Craiyon and OpenAI's DALL-E 2 text-to-image models and evaluated on the three criteria proposed for determining the quality of visual AIGC: (1) the number of objects, (2), resolution (strictly image quality; label/prompt exclusive), and (3) representativeness (consideration for how well the image matches the label/prompt). We observe that the paid, DALL-E 2 model, produced a dataset with fewer objects per image, higher resolution, and higher representativeness compared to Craiyon (free). There is an inverse relationship between the number of objects/images and its resolution and representativeness. This study establishes three subjective metrics for the evaluation of synthetic images to support the creation of more inclusive AIGC.

Cite

CITATION STYLE

APA

Combs, K., Bihl, T. J., Gadre, A., & Christopherson, I. (2024). A Human-factors Approach for Evaluating AI-generated Images. In SIGMIS-CPR 2024 - Proceedings of the Computers and People Research Conference: Trust and Legitimacy in Emerging Technologies: Organizational and Societal Implications for People, Places and Power. Association for Computing Machinery, Inc. https://doi.org/10.1145/3632634.3655849

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free