A statistical approach for multilingual document clustering and topic extraction from clusters

  • Silva J
  • Mexia J
  • Coelho C
  • et al.
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

This paper describes a statistics-based methodology for document unsupervised clustering and cluster topics extraction. For this purpose, multiword lexical units (MWUs) of any length are automatically extracted from corpora using the LiPXtractor - a language independent statistics-based tool. The MWUs are taken as base features to characterize documents. These features are transformed and a document similarity matrix is constructed. From this matrix, a reduced set of features is selected using an approach based on Principal Component Analisys. Then, using the Model Based Clustering Analisys software, it is possible to obtain the best number of clusters. Precision and Recall for document-cluster assignment range above 90 %. Most important MWUs are extracted from each cluster and taken as document cluster topics. Results on new document classification will just be mentioned.

Cite

CITATION STYLE

APA

Silva, J., Mexia, J., Coelho, C. A., & Lopes, G. (2004). A statistical approach for multilingual document clustering and topic extraction from clusters. Pliska Studia Mathematica Bulgarica, 16, 207–228. Retrieved from http://www.math.bas.bg/~pliska/

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free