Abstract
An open issue in the sentiment classification of texts written in Serbian is the effect of different forms of morphological normalization and the usefulness of leveraging large amounts of unlabeled texts. In this paper, we assess the impact of lemmatizers and stemmers for Serbian on classifiers trained and evaluated on the Serbian Movie Review Dataset. We also consider the effectiveness of using word embeddings, generated from a large unlabeled corpus, as classification features.
Author supplied keywords
Cite
CITATION STYLE
Batanović, V., & Nikolić, B. (2017). Sentiment classification of documents in Serbian: The effects of morphological normalization and word embeddings. Telfor Journal, 9(2), 104–109. https://doi.org/10.5937/telfor1702104b
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.