A Signal Processing Method for Text Language Identification

H. Hassanpour; M. M. AlyanNezhadi; M. Mohammadi

Journal ArticleOPEN ACCESS

A Signal Processing Method for Text Language Identification

International Journal of Engineering, Transactions A: Basics (2021) 34(6) 1413-1418

DOI: 10.5829/ije.2021.34.06c.04

5Citations

10Readers

Abstract

Language identification is a critical step prior to any natural language processing. In this paper, a signal processing method for Language Identification is proposed. Sequence of characters in a word and the order of words in stream identify the language. The sequence of characters in a stream provides a signature to recognize the language without understanding its meaning. The signature can be extracted using signal processing techniques via converting texts into time series. Although several research and commercial software have been developed to identify text language, they need a standard dictionary for each language. We proposed a dictionary independent method consisting of three main steps, I) preprocessing, II) clustering and finally III) classification .First, the texts are converted to time series using UTF-8 codes. Second, to group similar languages, the obtained series are clustered. Third, each cluster is decomposed into 32 sub-bands using a Wavelet packet, and 32 features are extracted from each sub-band. Also, a multilayer perceptron neural network is used to classify the extracted features. The proposed method was tested on our dataset with 31000 texts from 31 different languages. The proposed method achieved 72.20% accuracy for language identification.

Author supplied keywords

Language Identification Signal Processing Wavelet Packet Transform Artificial Neural Network

Cite

CITATION STYLE

APA

Hassanpour, H., AlyanNezhadi, M. M., & Mohammadi, M. (2021). A Signal Processing Method for Text Language Identification. International Journal of Engineering, Transactions A: Basics, 34(6), 1413–1418. https://doi.org/10.5829/ije.2021.34.06c.04

A Signal Processing Method for Text Language Identification

Abstract

Author supplied keywords

Cite

Register to see more suggestions