Concepts and Effectiveness of the Cover-Coefficient-Based Clustering Methodology for Text Databases

87Citations
Citations of this article
103Readers
Mendeley users who have this article in their library.

Abstract

A new algorithm for document clustering is introduced. The base concept of the algorithm, the cover coefficient 1990 concept, provides a means of estimating the number of clusters within a document database and related indexing and clustering analytically. The CC concept is used also to identify the cluster seeds and to form clusters with these seeds. It is shown that the complexity of the clustering process is very low. The retrieval experiments show that the information-retrieval effectiveness of the algorithm is compatible with a very demanding complete linkage clustering method that is known to have good retrieval performance. The experiments also show that the algorithm is 15.1 to 63.5 (with an average of 47.5) percent better than four other clustering algorithms in cluster-based information retrieval. The experiments have validated the indexing-clustering relationships and the complexity of the algorithm and have shown improvements in retrieval effectiveness. In the experiments two document databases are used: TODS214 and INSPEC. The latter is a common database with 12,684 documents. © 1990, ACM. All rights reserved.

Cite

CITATION STYLE

APA

Can, F., & Ozkarahan, E. A. (1990). Concepts and Effectiveness of the Cover-Coefficient-Based Clustering Methodology for Text Databases. ACM Transactions on Database Systems (TODS), 15(4), 483–517. https://doi.org/10.1145/99935.99938

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free