Abstract
A dialog system is required to take a user’s intimacy into account to converse comfortably, since the appropriate style and content of a conversation between familiar people can be different from those between strangers. This paper aims to estimate the level of intimacy of a speaker during a dialog, and mainly focuses on tackling the problem of the sparseness of the labeled datasets for this task. Since there is no publicly available dataset in Japanese, we construct a new dialog corpus annotated with speakers’ intimacy. In addition, we propose a novel semi-supervised learning method that uses both labeled and unlabeled data for the intimacy estimation task. The existing method, Simple instance-Adaptive self-Training (SAT), which iterates the assignment of pseudo-labels to unlabeled data and training a model, is extended so that external knowledge can be incorporated to enhance the quality of the pseudo-labels. Specifically, classification models of two other tasks, Dialog Act Classification and Emotion Recognition in Conversation, are used as external knowledge, since relatively good models of these tasks can be trained using existing large-scale datasets. The results of experiments show that the proposed method significantly outperforms baselines including SAT, especially when the amount of labeled data is small.
Author supplied keywords
Cite
CITATION STYLE
Miura, T., Kanai, H., Shirai, K., & Kertkeidkachorn, N. (2025). Semi-Supervised Learning for Estimation of a Speaker’s Intimacy Enhanced With Classification Models of Auxiliary Tasks. IEEE Access, 13, 36832–36844. https://doi.org/10.1109/ACCESS.2025.3544600
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.