When random sampling preserves privacy

Kamalika Chaudhuri; Nina Mishra

Conference ProceedingsOPEN ACCESS

When random sampling preserves privacy

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2006) 4117 LNCS 198-213

DOI: 10.1007/11818175_12

61Citations

67Readers

Abstract

Many organizations such as the U.S. Census publicly release samples of data that they collect about private citizens. These datasets are first anonymized using various techniques and then a small sample is released so as to enable "do-it-yourself" calculations. This paper investigates the privacy of the second step of this process: sampling. We observe that rare values - values that occur with low frequency in the table - can be problematic from a privacy perspective. To our knowledge, this is the first work that quantitatively examines the relationship between the number of rare values in a table and the privacy in a released random sample. If we require ∈-privacy (where the larger ∈ is, the worse the privacy guarantee) with probability at least 1 - δ, we say that a value is rare if it occurs in at most Õ(1/∈) rows of the table (ignoring log factors). If there are no rare values, then we establish a direct connection between sample size that is safe to release and privacy. Specifically, if we select each row of the table with probability at most ∈ then the sample is O(∈)-private with high probability. In the case that there are t rare values, then the sample is Õ(∈δ/t)- private with probability at least 1 - δ. © International Association for Cryptologic Research 2006.

Cite

CITATION STYLE

APA

Chaudhuri, K., & Mishra, N. (2006). When random sampling preserves privacy. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 4117 LNCS, pp. 198–213). Springer Verlag. https://doi.org/10.1007/11818175_12

When random sampling preserves privacy

Abstract

Cite

Register to see more suggestions