Abstract
Being able to reliably estimate self-disclosure - a key component of friendship and intimacy - from language is important for many psychology studies. We build single-task models on five self-disclosure corpora, but find that these models generalize poorly; the within-domain accuracy of predicted message-level self-disclosure of the best-performing single-task model (mean Pearson's r=0.69) is much higher than the respective across data set accuracy (mean Pearson's r=0.32), due to both variations in the corpora (e.g., medical vs. general topics) and labelling instructions (target variables: self-disclosure, emotional disclosure, intimacy). However, some lexical features, such as expression of negative emotions and use of first person personal pronouns such as'I' reliably predict self-disclosure across corpora. We develop a multi-task model that improves results, with an average Pearson's r of 0.37 for out-of-corpora prediction.
Cite
CITATION STYLE
Reuel, A. K., Peralta, S., Sedoc, J., Sherman, G., & Ungar, L. (2022). Measuring the Language of Self-Disclosure across Corpora. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 1035–1047). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-acl.83
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.