Abstract
This paper describes an approach for the classification of millions of existing multi-word entities (MWEntities), such as organisation or event names, into thirteen category types, based only on the tokens they contain. In order to classify our very large in-house collection of multilingual MWEntities into an application-oriented set of entity categories, we trained and tested distantly-supervised classifiers in 43 languages based on MWEntities extracted from BabelNet. The best-performing classifier was the multi-class SVM using a TF.IDF-weighted data representation. Interestingly, one unique classifier trained on a mix of all languages consistently performed better than classifiers trained for individual languages, reaching an averaged F1-value of 88.8%. In this paper, we present the training and test data, including a human evaluation of its accuracy, describe the methods used to train the classifiers, and discuss the results.
Cite
CITATION STYLE
Chesney, S., Jacquet, G., Steinberger, R., & Piskorski, J. (2017). Multi-word Entity Classification in a Highly Multilingual Environment. In MWE 2017 - 13th Workshop on Multiword Expressions, Proceedings of the Workshop (pp. 11–20). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w17-1702
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.