Abstract
Although advanced text mining methods specifically adapted to the biomedical domain are continuously being developed, their applications on large scale have been scarce. One of the main reasons for this is the lack of computational resources and workforce required for processing large text corpora. In this paper we present a publicly available resource distributing preprocessed biomedical literature including sentence splitting, tokenization, part-of-speech tagging, syntactic parses and named entity recognition. The aim of this work is to support the future development of largescale text mining resources by eliminating the time consuming but necessary preprocessing steps. This resource covers the whole of PubMed and PubMed Central Open Access section, currently containing 26M abstracts and 1.4M full articles, constituting over 388M analyzed sentences. The resource is based on a fully automated pipeline, guaranteeing that the distributed data is always up-to-date. The resource is available at https://turkunlp. github.io/pubmed_parses/.
Cite
CITATION STYLE
Hakala, K., Kaewphan, S., Salakoski, T., & Ginter, F. (2016). Syntactic analyses and named entity recognition for pubmed and pubmed central-up-to-the-minute. In BioNLP 2016 - Proceedings of the 15th Workshop on Biomedical Natural Language Processing (pp. 102–107). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w16-2913
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.