Document clustering - Concepts, metrics and algorithms

9Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

Abstract

Document clustering, which is also refered to as text clustering, is a technique of unsupervised document organisation. Text clustering is used to group documents into subsets that consist of texts that are similar to each orher. These subsets are called clusters. Document clustering algorithms are widely used in web searching engines to produce results relevant to a query. An example of practical use of those techniques are Yahoo! hierarchies of documents [1]. Another application of document clustering is browsing which is defined as searching session without well specific goal. The browsing techniques heavily relies on document clustering. In this article we examine the most important concepts related to document clustering. Besides the algorithms we present comprehensive discussion about representation of documents, calculation of similarity between documents and evaluation of clusters quality.

Cite

CITATION STYLE

APA

Tarczynski, T. (2011). Document clustering - Concepts, metrics and algorithms. International Journal of Electronics and Telecommunications, 57(3), 271–277. https://doi.org/10.2478/v10177-011-0036-5

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free