Abstract
We work with Algerian, an under-resourced non-standardised Arabic variety, for which we compile a new parallel corpus consisting of user-generated textual data matched with normalised and corrected human annotations following data-driven and our linguistically motivated standard. We use an end-toend deep neural model designed to deal with context-dependent spelling correction and normalisation. Results indicate that a model with two CNN sub-network encoders and an LSTM decoder performs the best, and that word context matters. Additionally, preprocessing data token-by-token with an editdistance based aligner significantly improves the performance. We get promising results for the spelling correction and normalisation, as a pre-processing step for downstream tasks, on detecting binary Semantic Textual Similarity.
Cite
CITATION STYLE
Adouane, W., Bernardy, J. P., & Dobnik, S. (2019). Normalising non-standardised orthography in algerian code-switched user-generated data. In W-NUT@EMNLP 2019 - 5th Workshop on Noisy User-Generated Text, Proceedings (pp. 131–140). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/d19-5518
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.