NERDz: A Preliminary Dataset of Named Entities for Algerian

9Citations
Citations of this article
24Readers
Mendeley users who have this article in their library.
Get full text

Abstract

This paper introduces a first step towards creating the NERDz dataset. A manually annotated dataset of named entities for the Algerian vernacular dialect. The annotations are built on top of a recent extension to the Algerian NArabizi Treebank, comprizing NArabizi sentences with manual transliterations into Arabic and code-switched scripts. NERDz is therefore not only the first dataset of named entities for Algerian, but it also comprises parallel entities written in Latin, Arabic, and code-switched scripts. We present a detailed overview of our annotations, inter-annotator agreement measures, and define two preliminary baselines using a neural sequence labeling approach and an Algerian BERT model. We also make the annotation guidelines and the annotations available for future work.

Cite

CITATION STYLE

APA

Touileb, S. (2022). NERDz: A Preliminary Dataset of Named Entities for Algerian. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing: Long Paper, AACL-IJCNLP 2022 (Vol. 3, pp. 95–101). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.aacl-short.13

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free