Exploring Straightforward Methods for Automatic Conversational Red-Teaming

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Large language models (LLMs) are increasingly used in business dialogue systems but they also pose security and ethical risks. Multi-turn conversations, in which context influences the model's behavior, can be exploited to generate undesired responses. In this paper, we investigate the use of off-the-shelf LLMs in conversational red-teaming settings, where an attacker LLM attempts to elicit undesired outputs from a target LLM. Our experiments address critical questions and offer valuable insights regarding the effectiveness of using LLMs as automated red-teamers, shedding light on key strategies and usage approaches that significantly impact their performance. Our findings demonstrate that off-the-shelf models can serve as effective red-teamers, capable of adapting their attack strategies based on prior attempts. Allowing these models to freely steer conversations and conceal their malicious intent further increases attack success. However, their effectiveness decreases as the alignment of the target model improves.

Cite

CITATION STYLE

APA

Kour, G., Zwerdling, N., Zalmanovici, M., Anaby-Tavor, A., Fandina, O. N., & Farchi, E. (2025). Exploring Straightforward Methods for Automatic Conversational Red-Teaming. In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 3, pp. 112–128). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-industry.10

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free