Abstract
As AI systems increasingly shape decision-making in creative design contexts, understanding how humans engage with these tools has become a critical challenge for interactive intelligent systems research. This article contributes a challenge to rethink how to evaluate human–AI collaborative systems, advocating for a more nuanced and multidimensional approach. Findings from one of the largest field studies to date (n = 808) of a human–AI co-creative system, The Genetic Car Designer, complemented by a controlled lab study (n = 12), are presented. The system is based on an interactive evolutionary algorithm where participants were tasked with designing a simple 2D representation of a car. Participants were exposed to galleries of design suggestions generated by an intelligent system, MAP–Elites, and a random control. Results indicate that exposure to galleries generated by MAP–Elites significantly enhanced both cognitive and behavioral engagement, leading to higher-quality design outcomes. Crucially for the wider community, the analysis reveals that conventional evaluation methods, which often focus on solely behavioral and design quality metrics, fail to capture the full spectrum of user engagement. By considering the human–AI design process as a changing emotional, behavioral, and cognitive state of the designer, we propose evaluating human–AI systems holistically and considering intelligent systems as a core part of the user experience—not simply a back-end tool.
Author supplied keywords
Cite
CITATION STYLE
Walton, S. P., Evans, B. J., Rahat, A. A. M., Stovold, J., & Vincalek, J. (2026). From Metrics to Meaning: Time to Rethink Evaluation in Human–AI Collaborative Design. ACM Transactions on Interactive Intelligent Systems, 16(1). https://doi.org/10.1145/3773292
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.