Scientific documents clustering based on text summarization

12Citations
Citations of this article
21Readers
Mendeley users who have this article in their library.

Abstract

In this paper a novel method is proposed for scientific document clustering. The proposed method is a summarization-based hybrid algorithm which comprises a preprocessing phase. In the summarization phase unimportant words which are not frequently used in the document are removed. This process reduces the amount of data for the clustering purpose. In this proposed method after the preprocessing phase, Term Frequency/Inverse Document Frequency (TFIDF) is calculated for all words in the document and BM25 in caluculated for words in sentences and summed over the document to score each word in document level. In next phase, Text summarization is performed based on BM25 scores. After that document clustering is done according to the scores of calculated TFIDF. The hybrid progress of the proposed scheme, from preprocessing phase to cluster labeling, gains a rapid and efficient clustering method which is evaluated by 400 English texts extracted from scientific articles of 11 different topics. The proposed method is compared with CSSA, SMTC and Max-Capture methods. The results demonstrate the proficiency of the proposed scheme in terms of computation time, and comparative efficiency using F-measure criterion.

Cite

CITATION STYLE

APA

Amoli, P. V., & Sojoodi Sh., O. (2015). Scientific documents clustering based on text summarization. International Journal of Electrical and Computer Engineering, 5(4), 782–787. https://doi.org/10.11591/ijece.v5i4.pp782-787

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free