Cformer: Semi-Supervised Text Clustering Based on Pseudo Labeling

7Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We propose a semi-supervised learning method called Cformer for automatic clustering of text documents in cases where clusters are described by a small number of labeled examples, while the majority of training examples are unlabeled. We motivate this setting with an application in contextual programmatic advertising, a type of content placement on news pages that does not exploit personal information about visitors but relies on the availability of a high-quality clustering computed on the basis of a small number of labeled samples. To enable text clustering with little training data, Cformer leverages the teacher-student architecture of Meta Pseudo Labels. In addition to unlabeled data, Cformer uses a small amount of labeled data to describe the clusters aimed at. Our experimental results confirm that the performance of the proposed model improves the state-of-the-art if a reasonable amount of labeled data is available. The models are comparatively small and suitable for deployment in constrained environments with limited computing resources. The source code is available at https://github.com/Aha6988/Cformer.

Cite

CITATION STYLE

APA

Hatefi, A., Vu, X. S., Bhuyan, M., & Drewes, F. (2021). Cformer: Semi-Supervised Text Clustering Based on Pseudo Labeling. In International Conference on Information and Knowledge Management, Proceedings (pp. 3078–3082). Association for Computing Machinery. https://doi.org/10.1145/3459637.3482073

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free