Constructing linguistic resources for the tunisian dialect using textual user-generated contentson the social web

Jihen Younes; Hadhemi Achour; Emna Souissi

Conference ProceedingsOPEN ACCESS

Constructing linguistic resources for the tunisian dialect using textual user-generated contentson the social web

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2015) 9396 3-14

DOI: 10.1007/978-3-319-24800-4_1

19Citations

13Readers

Abstract

In Arab countries, the dialect is daily gaining ground in the social interaction on the web and swiftly adapting to globalization. Strengthening the relationship of its practitioners with the outside world and facilitating their social exchanges, the dialect encompasses every day new transcriptions that arouse the curiosity of researchers in the NLP community. In this article, we focus specifically on the Tunisian dialect processing. Our goal is to build corpora and dictionaries allowing us to begin our study of this language and to identify its specificities. As a first step, we extract textual user-generated contents on the social Web, we then conduct an automatic content filtering and classification, leaving only the texts containing Tunisian dialect. Finally, we present some of its salient features from the built corpora.

Author supplied keywords

Cite

CITATION STYLE

APA

Younes, J., Achour, H., & Souissi, E. (2015). Constructing linguistic resources for the tunisian dialect using textual user-generated contentson the social web. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 9396, pp. 3–14). Springer Verlag. https://doi.org/10.1007/978-3-319-24800-4_1

Constructing linguistic resources for the tunisian dialect using textual user-generated contentson the social web

Abstract

Author supplied keywords

Cite

Register to see more suggestions