Abstract
In this paper, we study the information lost when a real-valued statistic is used to reduce or summarize sample data from a discrete random variable with a one-dimensional parameter. We compare the probability that a random sample gives a particular data set to the probability of the statistic’s value for this data set. We focus on sufficient statistics for the parameter of interest and develop a general formula independent of the parameter for the Shannon information lost when a data sample is reduced to such a summary statistic. We also develop a measure of entropy for this lost information that depends only on the real-valued statistic but neither the parameter nor the data. Our approach would also work for non-sufficient statistics, but the lost information and associated entropy would involve the parameter. The method is applied to three well-known discrete distributions to illustrate its implementation.
Author supplied keywords
Cite
CITATION STYLE
Moghimi, M., & Corley, H. W. (2020). Information loss due to the data reduction of sample data from discrete distributions. Data, 5(3), 1–18. https://doi.org/10.3390/data5030084
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.