Abstract
Nowcasting based on social media text promises to provide unobtrusive and near real-time predictions of community-level outcomes. These outcomes are typically regarding people, but the data is often aggregated without regard to users in the Twitter populations of each community. This paper describes a simple yet effective method for building community-level models using Twitter language aggregated by user. Results on four different U.S. county-level tasks, spanning demographic, health, and psychological outcomes show large and consistent improvements in prediction accuracies (e.g. from Pearson r = .73 to .82 for median income prediction or r = .37 to .47 for life satisfaction prediction) over the standard approach of aggregating all tweets. We make our aggregated and anonymized community-level data, derived from 37 billion tweets - over 1 billion of which were mapped to counties, available for research.
Cite
CITATION STYLE
Giorgi, S., Preoţiuc-Pietro, D., Buffone, A., Rieman, D., Ungar, L. H., & Andrew Schwartz, H. (2018). The remarkable benefit of user-level aggregation for lexical-based population-level predictions. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018 (pp. 1167–1172). Association for Computational Linguistics. https://doi.org/10.18653/v1/d18-1148
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.