Multilingual document clustering using wikipedia as external knowledge

Kiran Kumar N.; G. S.K. Santosh; Vasudeva Varma

Conference Proceedings

Multilingual document clustering using wikipedia as external knowledge

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2011) 6653 LNCS 108-117

DOI: 10.1007/978-3-642-21353-3_9

10Citations

10Readers

Get full text

Abstract

This paper presents Multilingual Document Clustering (MDC) on comparable corpora. Wikipedia has evolved to be a major structured multilingual knowledge base. It has been highly exploited in many monolingual clustering approaches and also in comparing multilingual corpora. But there is no prior work which studied the impact of Wikipedia on MDC. Here, we have studied availing Wikipedia in enhancing MDC performance. We have leveraged Wikipedia knowledge structure (such as cross-lingual links, category, outlinks, Infobox information, etc.) to enrich the document representation for clustering multilingual documents. We have implemented Bisecting k-means clustering algorithm and experiments are conducted on a standard dataset provided by FIRE for their 2010 Ad-hoc Cross-Lingual document retrieval task on Indian languages. We have considered English and Hindi datasets for our experiments. By avoiding language-specific tools, our approach provides a general framework which can be easily extendable to other languages. The system was evaluated using F-score and Purity measures and the results obtained were encouraging. © 2011 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Kumar N., K., Santosh, G. S. K., & Varma, V. (2011). Multilingual document clustering using wikipedia as external knowledge. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 6653 LNCS, pp. 108–117). https://doi.org/10.1007/978-3-642-21353-3_9

Multilingual document clustering using wikipedia as external knowledge

Abstract

Author supplied keywords

Cite

Register to see more suggestions