SUAR: Towards Building a Corpus for the Saudi Dialect

33Citations
Citations of this article
62Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This paper presents the preliminary results of the construction of a morphologically annotated corpus for the Saudi dialect. We call the corpus SUAR (SaUdi corpus for NLP Applications and Resources). The corpus consists of around 104,079 words collected from different online sources. The linguistic features of the Saudi dialect are elaborated and compared with Modern Standard Arabic and other Arabic dialects. This paper conducts a pilot study to explore possible directions to facilitate the morphological annotation of the Saudi corpus. The corpus was automatically annotated using the MADAMIRA tool, after which it was manually inspected to validate the resulting analysis.

Cite

CITATION STYLE

APA

Al-Twairesh, N., Al-Matham, R., Madi, N., Almugren, N., Al-Aljmi, A. H., Alshalan, S., … Alfutamani, A. (2018). SUAR: Towards Building a Corpus for the Saudi Dialect. In Procedia Computer Science (Vol. 142, pp. 72–82). Elsevier B.V. https://doi.org/10.1016/j.procs.2018.10.462

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free