Abstract
In the current era of global internet connectivity, privacy concerns are of the utmost importance. When official statistical agencies collect spatially referenced, confidential data that they intend to release as public-use files, the suppression of small counts is a common measure that agencies take to protect the confidentiality of the datasubjects from ill-intentioned users. The goal of this paper is to demonstrate that an interval suppression criterion that does not suppress zeros can fail to protect regions with a single occurrence. We illustrate the difference in disclosure risk between an interval suppression criterion and a one-sided suppression criterion by considering a US county-level dataset composed of the number of deaths due to stroke in White men. Here, we illustrate that an interval suppression criterion leads to a twofold increase in the disclosure risk when compared with a one-sided suppression criterion for regions with a single incidence among a population of less than 600. We conclude with an extension of these findings beyond stroke mortality and by offering general guidelines for data suppression.
Author supplied keywords
Cite
CITATION STYLE
Quick, H., Holan, S. H., & Wikle, C. K. (2015). Zeros and ones: A case for suppressing zeros in sensitive count data with an application to stroke mortality. Stat, 4(1), 227–234. https://doi.org/10.1002/sta4.92
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.