Abstract
We report the results of a study into the use of a linear interpolating hidden Markov mo del (HMM) for the task of extracting technical terminology from MEDLINE abstracts and texts in the molecular-biology domain. This is the first stage in a system that will extract event information for automatically up dating biology databases. We trained the HMM entirely with bigrams based on lexical and character features in a relatively small corpus of 100 MEDLINE abstracts that were marked-up by domain experts with term classes such as proteins and DNA. Using cross-validation methods we achieved an F-score of 0.73 and we examine the contribution made by each part of the interpolation model to overcoming data sparseness.
Cite
CITATION STYLE
Collier, N., Nobata, C., & Tsujii, J. I. (2000). Extracting the Names of Genes and Gene Products with a Hidden Markov Model. In 18th International Conference on Computational Linguistics, COLING 2000 - Proceedings of Science (Vol. 1, pp. 201–207). Association for Computational Linguistics (ACL). https://doi.org/10.3115/990820.990850
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.