Abstract
We propose a semi-supervised learning method called Cformer for automatic clustering of text documents in cases where clusters are described by a small number of labeled examples, while the majority of training examples are unlabeled. We motivate this setting with an application in contextual programmatic advertising, a type of content placement on news pages that does not exploit personal information about visitors but relies on the availability of a high-quality clustering computed on the basis of a small number of labeled samples. To enable text clustering with little training data, Cformer leverages the teacher-student architecture of Meta Pseudo Labels. In addition to unlabeled data, Cformer uses a small amount of labeled data to describe the clusters aimed at. Our experimental results confirm that the performance of the proposed model improves the state-of-the-art if a reasonable amount of labeled data is available. The models are comparatively small and suitable for deployment in constrained environments with limited computing resources. The source code is available at https://github.com/Aha6988/Cformer.
Author supplied keywords
Cite
CITATION STYLE
Hatefi, A., Vu, X. S., Bhuyan, M., & Drewes, F. (2021). Cformer: Semi-Supervised Text Clustering Based on Pseudo Labeling. In International Conference on Information and Knowledge Management, Proceedings (pp. 3078–3082). Association for Computing Machinery. https://doi.org/10.1145/3459637.3482073
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.