Document clustering, which is also refered to as text clustering, is a technique of unsupervised document organisation. Text clustering is used to group documents into subsets that consist of texts that are similar to each orher. These subsets are called clusters. Document clustering algorithms are widely used in web searching engines to produce results relevant to a query. An example of practical use of those techniques are Yahoo! hierarchies of documents [1]. Another application of document clustering is browsing which is defined as searching session without well specific goal. The browsing techniques heavily relies on document clustering. In this article we examine the most important concepts related to document clustering. Besides the algorithms we present comprehensive discussion about representation of documents, calculation of similarity between documents and evaluation of clusters quality.
CITATION STYLE
Tarczynski, T. (2011). Document clustering - Concepts, metrics and algorithms. International Journal of Electronics and Telecommunications, 57(3), 271–277. https://doi.org/10.2478/v10177-011-0036-5
Mendeley helps you to discover research relevant for your work.