Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets

97Citations
Citations of this article
63Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

We introduce a new class of attacks on machine learning models. We show that an adversary who can poison a training dataset can cause models trained on this dataset to leak significant private details of training points belonging to other parties. Our active inference attacks connect two independent lines of work targeting the integrity and privacy of machine learning training data. Our attacks are effective across membership inference, attribute inference, and data extraction. For example, our targeted attacks can poison <0.1% of the training dataset to boost the performance of inference attacks by 1 to 2 orders of magnitude. Further, an adversary who controls a significant fraction of the training data (e.g., 50%) can launch untargeted attacks that enable 8× more precise inference on all other users' otherwise-private data points. Our results cast doubts on the relevance of cryptographic privacy guarantees in multiparty computation protocols for machine learning, if parties can arbitrarily select their share of training data.

Cite

CITATION STYLE

APA

Tramèr, F., Shokri, R., San Joaquin, A., Le, H., Jagielski, M., Hong, S., & Carlini, N. (2022). Truth Serum: Poisoning Machine Learning Models to Reveal Their Secrets. In Proceedings of the ACM Conference on Computer and Communications Security (pp. 2779–2792). Association for Computing Machinery. https://doi.org/10.1145/3548606.3560554

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free