Data Poisoning for In-context Learning

5Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.
Get full text

Abstract

In-context learning (ICL) has emerged as a capability of large language models (LLMs), enabling them to adapt to new tasks using provided examples. While ICL has demonstrated its strong effectiveness, there is limited understanding of its vulnerability against potential threats. This paper examines ICL’s vulnerability to data poisoning attacks. We introduce ICLPoison, an attacking method specially designed to exploit ICL’s unique learning mechanisms by identifying discrete text perturbations that influence LLM hidden states. We propose three representative attack strategies, evaluated across various models and tasks. Our experiments, including those on GPT-4, show that ICL performance can be significantly compromised by these attacks, highlighting the urgent need for improved defense mechanisms to protect LLMs’ integrity and reliability.

Cite

CITATION STYLE

APA

He, P., Xu, H., Xing, Y., Liu, H., Yamada, M., & Tang, J. (2025). Data Poisoning for In-context Learning. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 1680–1700). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.91

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free