Measuring the Language of Self-Disclosure across Corpora

16Citations
Citations of this article
45Readers
Mendeley users who have this article in their library.

Abstract

Being able to reliably estimate self-disclosure - a key component of friendship and intimacy - from language is important for many psychology studies. We build single-task models on five self-disclosure corpora, but find that these models generalize poorly; the within-domain accuracy of predicted message-level self-disclosure of the best-performing single-task model (mean Pearson's r=0.69) is much higher than the respective across data set accuracy (mean Pearson's r=0.32), due to both variations in the corpora (e.g., medical vs. general topics) and labelling instructions (target variables: self-disclosure, emotional disclosure, intimacy). However, some lexical features, such as expression of negative emotions and use of first person personal pronouns such as'I' reliably predict self-disclosure across corpora. We develop a multi-task model that improves results, with an average Pearson's r of 0.37 for out-of-corpora prediction.

Cite

CITATION STYLE

APA

Reuel, A. K., Peralta, S., Sedoc, J., Sherman, G., & Ungar, L. (2022). Measuring the Language of Self-Disclosure across Corpora. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 1035–1047). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-acl.83

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free