How Not to Detect Prompt Injections with an LLM

2Citations
Citations of this article
19Readers
Mendeley users who have this article in their library.
Get full text

Abstract

LLM-integrated applications and agents are vulnerable to prompt injection attacks, where adversaries embed malicious instructions within seemingly benign input data to manipulate the LLM's intended behavior. Recent defenses based on known-answer detection (KAD) scheme have reported near-perfect performance by observing an LLM's output to classify input data as clean or contaminated. KAD attempts to repurpose the very susceptibility to prompt injection as a defensive mechanism. We formally characterize the KAD scheme and uncover a structural vulnerability that invalidates its core security premise. To exploit this fundamental vulnerability, we methodically design an adaptive attack, DataFlip. It consistently evades KAD defenses, achieving detection rates as low as while reliably inducing malicious behavior with a success rate of - all without requiring white-box access to the LLM or any optimization procedures. We release our evaluation code at [10].

Cite

CITATION STYLE

APA

Choudhary, S., Anshumaan, D., Palumbo, N., & Jha, S. (2025). How Not to Detect Prompt Injections with an LLM. In Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025 (pp. 218–229). Association for Computing Machinery, Inc. https://doi.org/10.1145/3733799.3762980

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free