A New Feature Extraction, Reduction, and Classification Method for Documents Based on Fourier Transformation

1Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Text classification is the automated technique used to classify text into a predefined category that is more related to text. Most studies and researches have focused on classification texts written in English rather than Arabic because of the Arabic nature and difficulty of its structures. The difficult nature of Arabic makes it more complex and difficult to deal with because of its many rules and characteristics that are unique to it, but it has become necessary to deal with this language because of its widespread use on the internet. Feature representation and extraction, especially for text, have attracted considerable attention in recent years. The main object of this process is to convert text into a numerical representation. All approaches and systems depend on the frequency of words within text, but these approaches are few. This paper presents a new feature extraction method aimed at furthering natural language processing applications in any language, especially Arabic. It is based on transforming a vector of representation in the time domain into the frequency domain. Fourier feature extraction is a powerful and versatile technique. The primary goal of this approach is to extract salient features from raw text, and then a filter is used to remove noise from the extracted features. Its scalability and efficiency make it suitable to be used on large datasets, making it a widely adopted feature extraction technique. For classification, we used logistic regression, which is a powerful tool for classifying text data and offers significant advantages such as speed and accuracy. We used three world datasets (CNN, OSAC, and SANAD), and the results showed a good performance. The proposed method outperformed recent approaches, in which accuracy reached 98%.

Cite

CITATION STYLE

APA

Alfartosy, H. H., & Khafaji, H. K. (2023). A New Feature Extraction, Reduction, and Classification Method for Documents Based on Fourier Transformation. International Journal of Intelligent Engineering and Systems, 16(5), 586–597. https://doi.org/10.22266/ijies2023.1031.50

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free