Abstract
The value and nature of the representations learned during the pretraining of genomic language models (gLMs) remain actively debated. We introduce Nucleotide Generative Pretrained Transformer (GPT), a decoder-only transformer with single-nucleotide tokenization, to dissect the role of pretraining. Through experiments varying repetitive element (RE) weights during pretraining (0.0–1.0), comparative finetuning against random initialization, linear probing of internal representations, and sparse autoencoder (SAE)-based interpretability, we evaluated the impact of pretraining and how REs in genomic data influence model learning. Models with moderate RE downweighting (0.5) consistently achieved optimal performance across seven genomic classification tasks, with pretrained models providing substantial performance gains over baselines. SAE feature annotation via sequence alignment revealed substantial RE-associated patterns in the pretrained model internal representations, suggesting that REs—which comprise 30%–60% of mammalian genomes—may dominate the pretraining objective. Our findings support the utility of pretraining and underscore the need for pretraining strategies that better accommodate repetitive sequences across the genome while also fostering the learning of less common but biologically important representations. This study highlights a key challenge for gLMs: ensuring that models broadly learn functional genomic syntax beyond simply recognizing ubiquitous repeats.
Author supplied keywords
Cite
CITATION STYLE
Mclaughlin, S. M., & Lim, D. A. (2026). Probing genomic language models: Nucleotide Generative Pretrained Transformer and the role of pretraining in learned representations. Briefings in Bioinformatics, 27(1). https://doi.org/10.1093/bib/bbag011
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.