Abstract
Word embedding in the NLP area has attracted increasing attention in recent years. The continuous bag-of-words model (CBOW) and the continuous Skip-gram model (Skip-gram) have been developed to learn distributed representations of words from a large amount of unlabeled text data. In this paper, we explore the idea of integrating extra knowledge to the CBOW and Skip-gram models and applying the new models to biomedical NLP tasks. The main idea is to construct a weighted graph from knowledge bases (KBs) to represent structured relationships among words/concepts. In particular, we propose a GCBOW model and a GSkip-gram model respectively by integrating such a graph into the original CBOW model and Skip-gram model via graph regularization. Our experiments on four general domain standard datasets show encouraging improvements with the new models. Further evaluations on two biomedical NLP tasks (biomedical similarity/relatedness task and biomedical Information Retrieval (IR) task) show that our methods have better performance than baselines.
Cite
CITATION STYLE
Ling, Y., An, Y., Liu, M., Hasan, S. A., Fan, Y., & Hu, X. (2017). Integrating extra knowledge into word embedding models for biomedical NLP tasks. In Proceedings of the International Joint Conference on Neural Networks (Vol. 2017-May, pp. 968–975). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/IJCNN.2017.7965957
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.