Abstract
In this paper a novel method is proposed for scientific document clustering. The proposed method is a summarization-based hybrid algorithm which comprises a preprocessing phase. In the summarization phase unimportant words which are not frequently used in the document are removed. This process reduces the amount of data for the clustering purpose. In this proposed method after the preprocessing phase, Term Frequency/Inverse Document Frequency (TFIDF) is calculated for all words in the document and BM25 in caluculated for words in sentences and summed over the document to score each word in document level. In next phase, Text summarization is performed based on BM25 scores. After that document clustering is done according to the scores of calculated TFIDF. The hybrid progress of the proposed scheme, from preprocessing phase to cluster labeling, gains a rapid and efficient clustering method which is evaluated by 400 English texts extracted from scientific articles of 11 different topics. The proposed method is compared with CSSA, SMTC and Max-Capture methods. The results demonstrate the proficiency of the proposed scheme in terms of computation time, and comparative efficiency using F-measure criterion.
Author supplied keywords
Cite
CITATION STYLE
Amoli, P. V., & Sojoodi Sh., O. (2015). Scientific documents clustering based on text summarization. International Journal of Electrical and Computer Engineering, 5(4), 782–787. https://doi.org/10.11591/ijece.v5i4.pp782-787
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.