UnibucKernel: A kernel-based learning method for complex word identification

Andrei M. Butnaru; Radu Tudor Ionescu

Conference ProceedingsOPEN ACCESS

UnibucKernel: A kernel-based learning method for complex word identification

Proceedings of the 13th Workshop on Innovative Use of NLP for Building Educational Applications, BEA 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018 (2018) 175-183

DOI: 10.18653/v1/w18-0519

8Citations

83Readers

Abstract

In this paper, we present a kernel-based learning approach for the 2018 Complex Word Identification (CWI) Shared Task. Our approach is based on combining multiple low-level features, such as character n-grams, with high-level semantic features that are either automatically learned using word embeddings or extracted from a lexical knowledge base, namely WordNet. After feature extraction, we employ a kernel method for the learning phase. The feature matrix is first transformed into a normalized kernel matrix. For the binary classification task (simple versus complex), we employ Support Vector Machines. For the regression task, in which we have to predict the complexity level of a word (a word is more complex if it is labeled as complex by more annotators), we employ ν-Support Vector Regression. We applied our approach only on the three English data sets containing documents from Wikipedia, WikiNews and News domains. Our best result during the competition was the third place on the English Wikipedia data set. However, in this paper, we also report better post-competition results.

Cite

CITATION STYLE

APA

Butnaru, A. M., & Ionescu, R. T. (2018). UnibucKernel: A kernel-based learning method for complex word identification. In Proceedings of the 13th Workshop on Innovative Use of NLP for Building Educational Applications, BEA 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018 (pp. 175–183). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-0519

UnibucKernel: A kernel-based learning method for complex word identification

Abstract

Cite

Register to see more suggestions