Abstract
Phenotype definitions are crucial for the progression of precision and personalized medicine. Although phenotype knowledge bases such as PheKB and the OHDSI library are available, they rely heavily on manual input. This study introduces PheCatcher, an automated pipeline that integrates BiomedBERT-based Named Entity Recognition (NER) and Relation Extraction (RE) to extract phenotypes and standardized codes from biomedical literature. To complement human annotation, GPT-4 was utilized to generate synthetic data, which improved model performance. The NER model's F1 score for "phenotype" entities increased from 0.616 to 0.800, and the RE model achieved an F1 score of 0.901. The application of the pipeline to the PubMed Central (PMC) articles resulted in the extraction of 173,283 phenotype definitions, which are now publicly accessible. Our study underscores the potential of synthetic data for information extraction (IE) and offers the first evidence of the feasibility of leveraging synthetic data to build a complete IE system.
Author supplied keywords
Cite
CITATION STYLE
Hu, Y., Hong, N., Li, Y., Peng, X., Chen, Y., & Xu, H. (2025). PheCatcher: Leveraging LLM-Generated Synthetic Data for Automated Phenotype Definition Extraction from Biomedical Literature. In Studies in Health Technology and Informatics (Vol. 329, pp. 718–722). IOS Press BV. https://doi.org/10.3233/SHTI250934
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.