Extracting the Names of Genes and Gene Products with a Hidden Markov Model

173Citations
Citations of this article
122Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We report the results of a study into the use of a linear interpolating hidden Markov mo del (HMM) for the task of extracting technical terminology from MEDLINE abstracts and texts in the molecular-biology domain. This is the first stage in a system that will extract event information for automatically up dating biology databases. We trained the HMM entirely with bigrams based on lexical and character features in a relatively small corpus of 100 MEDLINE abstracts that were marked-up by domain experts with term classes such as proteins and DNA. Using cross-validation methods we achieved an F-score of 0.73 and we examine the contribution made by each part of the interpolation model to overcoming data sparseness.

Cite

CITATION STYLE

APA

Collier, N., Nobata, C., & Tsujii, J. I. (2000). Extracting the Names of Genes and Gene Products with a Hidden Markov Model. In 18th International Conference on Computational Linguistics, COLING 2000 - Proceedings of Science (Vol. 1, pp. 201–207). Association for Computational Linguistics (ACL). https://doi.org/10.3115/990820.990850

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free