Describing the Pearson R distribution of aggregate data

5Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

Ecological studies and epidemiology need to use group averaged data to make inferences about individual patterns. However, using correlations based on averages to estimate correlations of individual scores is subject to an "ecological fallacy". The purpose of this article is to create distributions of Pearson R correlation values computed from grouped averaged or aggregate data using Monte Carlo simulations and random sampling. We show that, as the group size increases, the distributions can be approximated by a generalized hypergeometric distribution. The expectation of the constructed distribution slightly underestimates the individual Pearson R value, but the difference becomes smaller as the number of groups increases. The approximate normal distribution resulting from Fisher's transformation can be used to build confidence intervals to approximate the Pearson R value based on individual scores from the Pearson R value based on the aggregated scores.

Cite

CITATION STYLE

APA

Torres, D. J. (2020). Describing the Pearson R distribution of aggregate data. Monte Carlo Methods and Applications, 26(1), 17–32. https://doi.org/10.1515/mcma-2020-2054

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free