Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

LLM agents are increasingly deployed to plan, retrieve, and write with tools, yet evaluation still leans on static benchmarks and small human studies. We present the Agent-Testing Agent (ATA), a meta-agent that combines static code analysis, developer interrogation, literature mining, and persona-driven adversarial test generation whose difficulty adapts via judge feedback. Each dialogue is scored with an LLM-as-a-Judge (LAAJ) rubric and used to steer subsequent tests toward the agent’s weakest capabilities. On a travel planner and a Wikipedia writer, the ATA surfaces more diverse and severe failures than expert annotators while matching severity, and finishes in 20–30 minutes versus ten-annotator rounds that took days. Ablating code analysis and web search increases variance and miscalibration, underscoring the value of evidence-grounded test generation. The ATA outputs quantitative metrics and qualitative bug reports for developers. We release the full open-source implementation.

Cite

CITATION STYLE

APA

Komoravolu, S., & Mrini, K. (2026). Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents. In EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers) (Vol. 1, pp. 7199–7214). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2026.eacl-long.339

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free