Learning multiview embeddings of Twitter users

Adrian Benton; Raman Arora; Mark Dredze

Conference ProceedingsOPEN ACCESS

Learning multiview embeddings of Twitter users

54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Short Papers (2016) 14-19

DOI: 10.18653/v1/p16-2003

77Citations

171Readers

Abstract

Low-dimensional vector representations are widely used as stand-ins for the text of words, sentences, and entire documents. These embeddings are used to identify similar words or make predictions about documents. In this work, we consider embeddings for social media users and demonstrate that these can be used to identify users who behave similarly or to predict attributes of users. In order to capture information from all aspects of a user's online life, we take a multiview approach, applying a weighted variant of Generalized Canonical Correlation Analysis (GCCA) to a collection of over 100,000 Twitter users. We demonstrate the utility of these multiview embeddings on three downstream tasks: user engagement, friend selection, and demographic attribute prediction.

Cite

CITATION STYLE

APA

Benton, A., Arora, R., & Dredze, M. (2016). Learning multiview embeddings of Twitter users. In 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Short Papers (pp. 14–19). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p16-2003

Learning multiview embeddings of Twitter users

Abstract

Cite

Register to see more suggestions