Abstract
Short texts usually encounter the problem of data sparseness, as they do not provide sufficient term co-occurrence information. In this paper, we show how to mitigate the problem in short text classification through word embeddings. We assume that a short text document is a specific sample of one distribution in a Gaussian-Bayesian framework. Furthermore, a fast clustering algorithm is utilized to expand and enrich the context of short text in embedding space. This approach is compared with those based on the classical bag-of-words approaches and neural network based methods. Experimental results validate the effectiveness of the proposed method.
Author supplied keywords
Cite
CITATION STYLE
Ma, C., Zhao, Q., Pan, J., & Yan, Y. (2016). Short text classification based on distributional representations of words. In IEICE Transactions on Information and Systems (Vol. E99D, pp. 2562–2565). Maruzen Co., Ltd. https://doi.org/10.1587/transinf.2016SLL0006
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.