Abstract
Accurately learning from user data while providing quantifiable privacy guarantees provides an opportunity to build better ML models while maintaining user trust. This paper presents a formal approach to carrying out privacy preserving text perturbation using the notion of dχ -privacy designed to achieve geo-indistinguishability in location data. Our approach applies carefully calibrated noise to vector representation of words in a high dimension space as defined by word embedding models. We present a privacy proof that satisfies dχ -privacy where the privacy parameter ε provides guarantees with respect to a distance metric defined by the word embedding space. We demonstrate how ε can be selected by analyzing plausible deniability statistics backed up by large scale analysis on GloVe and fastText embeddings. We conduct privacy audit experiments against 2 baseline models and utility experiments on 3 datasets to demonstrate the tradeoff between privacy and utility for varying values of ε on different task types. Our results demonstrate practical utility (< 2% utility loss for training binary classifiers) while providing better privacy guarantees than baseline models.
Cite
CITATION STYLE
Feyisetan, O., Balle, B., Drake, T., & Diethe, T. (2020). Privacy- And utility-preserving textual analysis via calibrated multivariate perturbations. In WSDM 2020 - Proceedings of the 13th International Conference on Web Search and Data Mining (pp. 178–186). Association for Computing Machinery, Inc. https://doi.org/10.1145/3336191.3371856
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.