PheCatcher: Leveraging LLM-Generated Synthetic Data for Automated Phenotype Definition Extraction from Biomedical Literature

14Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Phenotype definitions are crucial for the progression of precision and personalized medicine. Although phenotype knowledge bases such as PheKB and the OHDSI library are available, they rely heavily on manual input. This study introduces PheCatcher, an automated pipeline that integrates BiomedBERT-based Named Entity Recognition (NER) and Relation Extraction (RE) to extract phenotypes and standardized codes from biomedical literature. To complement human annotation, GPT-4 was utilized to generate synthetic data, which improved model performance. The NER model's F1 score for "phenotype" entities increased from 0.616 to 0.800, and the RE model achieved an F1 score of 0.901. The application of the pipeline to the PubMed Central (PMC) articles resulted in the extraction of 173,283 phenotype definitions, which are now publicly accessible. Our study underscores the potential of synthetic data for information extraction (IE) and offers the first evidence of the feasibility of leveraging synthetic data to build a complete IE system.

Cite

CITATION STYLE

APA

Hu, Y., Hong, N., Li, Y., Peng, X., Chen, Y., & Xu, H. (2025). PheCatcher: Leveraging LLM-Generated Synthetic Data for Automated Phenotype Definition Extraction from Biomedical Literature. In Studies in Health Technology and Informatics (Vol. 329, pp. 718–722). IOS Press BV. https://doi.org/10.3233/SHTI250934

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free