Abstract
Pinpointing causal genes at genome-wide association study (GWAS) loci remains a major bottleneck. Existing literature-mining approaches are often limited in accuracy and scalability. We show that large language models (LLMs) can accurately prioritize likely causal genes at GWAS loci. We systematically evaluated several widely available general-purpose LLMs against benchmark datasets of high-confidence causal genes, including a unique set from 23 unpublished GWAS. Our results demonstrate that LLMs outperform or match current state-of-the-art methods and, crucially, exhibit robust performance on novel loci not previously linked to traits, underscoring their generalizability. Moreover, when integrated with existing methods, LLMs substantially enhance overall performance. This work establishes LLMs as an accurate, scalable, and broadly generalizable approach to accelerate causal gene identification in complex traits.
Cite
CITATION STYLE
Shringarpure, S. S., Wang, W., Karagounis, S., Wang, X., Reisetter, A. C., Auton, A., & Khan, A. A. (2026). Large language models identify causal genes in complex trait GWAS. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing, 31, 480–493. https://doi.org/10.1142/9789819824755_0034
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.