Abstract
The process of plagiarism detection is one of the challenges in revealing the originality of a document, especially in the fields of science and research. Natural language processing methods can recognize and determine the level of similarity between different documents. In this paper, we tackle the task of extrinsic plagiarism detection based on semantic and syntactic approaches. The objective is to identify segments of a document that show strong similarity with a group of reference documents dealing with the same topic. In this paper, we present our hybrid approach that implements semantic and syntactic features based on Latent Dirichlet Allocation (LDA) and Wu & Plamer algorithm. The proposed approach has been evaluated on a PAN13 public dataset with a total accuracy of 85%.
Author supplied keywords
Cite
CITATION STYLE
Nahar, K. M. O., Alshtaiwi, M., Alikhashashneh, E., Shatnawi, N., Al-Shannaq, M. A., Abual-Rub, M., & Bani-Ismail, B. (2024). Plagiarism Detection System by Semantic and Syntactic Analysis Based on Latent Dirichlet Allocation Algorithm. International Journal of Advances in Soft Computing and Its Applications, 16(1), 40–55. https://doi.org/10.15849/IJASCA.240330.03
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.