Tagger for Polish computer mediated communication texts

5Citations
Citations of this article
60Readers
Mendeley users who have this article in their library.
Get full text

Abstract

In this paper we present a morphosyntactic tagger dedicated to Computer-mediated Communication texts in Polish. Its construction is based on an expanded RNN-based neural network adapted to the work on noisy texts. Among several techniques, the tagger utilises fastText embedding vectors, sequential character embedding vectors, and Brown clustering for the coarse-grained representation of sentence structures. In addition a set of manually written rules was proposed for postprocessing. The system was trained to disambiguate descriptions of words in relation to Parts of Speech tags together with the full morphological information in terms of values for the different grammatical categories. We present also evaluation of several model variants on the gold standard annotated CMC data, comparison to the state-of-the-art taggers for Polish and error analysis. The proposed tagger shows significantly better results in this domain and demonstrates the viability of adaptation.

Cite

CITATION STYLE

APA

Walentynowicz, W., Piasecki, M., & Oleksy, M. (2019). Tagger for Polish computer mediated communication texts. In International Conference Recent Advances in Natural Language Processing, RANLP (Vol. 2019-September, pp. 1295–1303). Incoma Ltd. https://doi.org/10.26615/978-954-452-056-4_148

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free