Normalising non-standardised orthography in algerian code-switched user-generated data

10Citations
Citations of this article
68Readers
Mendeley users who have this article in their library.

Abstract

We work with Algerian, an under-resourced non-standardised Arabic variety, for which we compile a new parallel corpus consisting of user-generated textual data matched with normalised and corrected human annotations following data-driven and our linguistically motivated standard. We use an end-toend deep neural model designed to deal with context-dependent spelling correction and normalisation. Results indicate that a model with two CNN sub-network encoders and an LSTM decoder performs the best, and that word context matters. Additionally, preprocessing data token-by-token with an editdistance based aligner significantly improves the performance. We get promising results for the spelling correction and normalisation, as a pre-processing step for downstream tasks, on detecting binary Semantic Textual Similarity.

Cite

CITATION STYLE

APA

Adouane, W., Bernardy, J. P., & Dobnik, S. (2019). Normalising non-standardised orthography in algerian code-switched user-generated data. In W-NUT@EMNLP 2019 - 5th Workshop on Noisy User-Generated Text, Proceedings (pp. 131–140). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/d19-5518

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free