In this paper we present a rule- and lexicon-based system for the recognition of Named Entities (NE) in Serbian newspaper texts that was used to prepare a gold standard annotated with personal names. It was further used to prepare training sets for four different levels of annotation, which were further used to train two Named Entity Recognition (NER) systems: Stanford and spaCy. All obtained models, together with a rule- and lexicon-based system were evaluated on two sample texts: a part of the gold standard and an independent newspaper text of approximately the same size. The results show that rule- and lexicon-based system outperforms trained models in all four scenarios (measured by F1), while Stanford models have the highest recall. The produced models are incorporated into a Web platform NER&Beyond that provides various NE-related functions.
CITATION STYLE
Šandrih, B., Krstev, C., & Stanković, R. (2019). Development and evaluation of three named entity recognition systems for Serbian - The case of personal names. In International Conference Recent Advances in Natural Language Processing, RANLP (Vol. 2019-September, pp. 1060–1068). Incoma Ltd. https://doi.org/10.26615/978-954-452-056-4_122
Mendeley helps you to discover research relevant for your work.