Perhitungan Kemiripan Term Co-occurence Berdasarkan Cluster Dokumen Untuk Pengembangan Thesaurus Bahasa Arab

  • Yunianto D
  • Arifin A
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Generating automatically thesaurus is by calculating the similarity value term. To get the value of the similarity can be carried out with the co-occurence approach is to see the frequency of occurrence along these terms. The frequency of how much these terms occurence  on documents corpus. Each of the documents contained in the corpus have content or topics vary. So the terms that are in the document a specific topic will have a different context with the terms of the document with other topics. Therefore, this paper proposes a new method of measurement term similarity with co-occurence based on cluster of documents on the generate of Arabic thesaurus. The documents will be in the corpus clustering to group by the proximity of the content of the document. To get the term similarity value calculation clusterweight by leveraging the value of inverse class frequency of each term to an existing cluster. Thesaurus is formed by looking at the value of the calculation result of the similarity term. Thesaurus formed by the proposed method succeeded in improving inter-term relevance is evidenced by the experimental results have a precision value of 63,3%, amounting to 78,6% recall and F-measure by 50%.

Cite

CITATION STYLE

APA

Yunianto, D. R., & Arifin, A. Z. (2017). Perhitungan Kemiripan Term Co-occurence Berdasarkan Cluster Dokumen Untuk Pengembangan Thesaurus Bahasa Arab. JURNAL INFOTEL, 9(1), 64. https://doi.org/10.20895/infotel.v9i1.168

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free