Abstract
In-context learning (ICL) enables large language models (LLMs) to adapt to new tasks using only a few examples, without requiring fine-tuning. However, the new privacy and security risks brought about by this increasing capability have not received enough attention, and there is a lack of research on this issue. In this work, we propose a novel membership inference attack (MIA) method, termed Neighborhood Deviation Attack, specifically designed to evaluate the privacy risks of LLMs in ICL. Unlike traditional MIA methods, our approach does not require access to model parameters and instead relies solely on analyzing the model’s output behavior. We first generate neighborhood prefixes for target samples and use the LLM, conditioned on ICL examples, to complete the text. We then compute the deviation between the original and completed texts and infer membership based on these deviations. We conduct experiments on three datasets and three LLMs and further explore the influence of key hyperparameters on the method’s performance and their underlying reasons. Experimental results show that our method is significantly better than the comparative methods in terms of stability and achieves better accuracy in most cases. Furthermore, we discuss four potential defense strategies, including increasing the diversity of ICL examples and introducing controlled randomness in the inference process to reduce the risk of privacy leakage.
Author supplied keywords
Cite
CITATION STYLE
Hou, D., Yang, Z., Zheng, L., Jin, B., Xu, H., Li, Y., … Peng, K. (2025). Neighborhood Deviation Attack Against In-Context Learning. Applied Sciences (Switzerland), 15(8). https://doi.org/10.3390/app15084177
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.