Pichia-CLM: A language model–based codon optimization pipeline for Komagataella phaffii

1Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The preference in synonymous codon usage—the so-called codon usage bias (CUB)—is governed by several factors such as the host organism, context and function of the gene, and the position of the codon within the gene itself. We demonstrated that this mapping can be learned from the host’s genome using language models and subsequently applied for codon optimization of heterologous proteins expressed by the host. This pipeline called Pichia–Codon language model (Pichia-CLM) was applied to the industrial host organism, Komagataella phaffii. With this approach, production of heterologous proteins was enhanced up to threefold compared to their native sequences. Furthermore, Pichia-CLM consistently yielded constructs with enhanced productivity for proteins of varied complexity, compared to commercially available tools. Finally, we showed that Pichia-CLM generates sequences resembling the properties of codon usage found in the host’s intrinsic host cell proteins and learned features such as avoiding negative cis-regulatory and repeat elements based on patterns in the genome data. These results show the potential of language models to unbiasedly learn patterns and design robust sequences for improved protein production.

Cite

CITATION STYLE

APA

Narayanan, H., & Love, J. C. (2026). Pichia-CLM: A language model–based codon optimization pipeline for Komagataella phaffii. Proceedings of the National Academy of Sciences of the United States of America, 123(8). https://doi.org/10.1073/pnas.2522052123

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free