Atoxia: Red-teaming Large Language Models with Target Toxic Answers

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Despite the substantial advancements in artificial intelligence, large language models (LLMs) remain being challenged by generation safety. With adversarial jailbreaking prompts, one can effortlessly induce LLMs to output harmful content, causing unexpected negative social impacts. This vulnerability highlights the necessity for robust LLM red-teaming strategies to identify and mitigate such risks before large-scale application. To detect specific types of risks, we propose a novel red-teaming method that Attacks LLMs with Target Toxic Answers (Atoxia). Given a particular harmful answer, Atoxia generates a corresponding user query and a misleading answer opening to examine the internal defects of a given LLM. The proposed attacker is trained within a reinforcement learning scheme with the LLM outputting probability of the target answer as the reward. We verify the effectiveness of our method on various red-teaming benchmarks, such as AdvBench and HH-Harmless. The empirical results demonstrate that Atoxia can successfully detect safety risks in not only open-source models but also state-of-the-art black-box models such as GPT-4o.

Cite

CITATION STYLE

APA

Du, Y., Li, Z., Cheng, P., Wan, X., & Gao, A. (2025). Atoxia: Red-teaming Large Language Models with Target Toxic Answers. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 3251–3266). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.179

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free