Abstract
Proteins act as the terminal effectors of cellular function, encoding the phenotypic consequences of genomic and transcriptomic programs. Although transcriptomic profiles serve as accessible proxies, they remain incomplete surrogates for the proteomic landscape that ultimately defines cellular phenotypes. Current single-cell foundation models, however, are trained exclusively on transcriptomes, resulting in biased and partial characterizations of cellular states. To address this limitation, we introduce CAPTAIN, a multimodal foundational model pretrained on over four million single cells with concurrently measured transcriptomes and a curated repertoire of 382 surface proteins across diverse human and mouse tissues. Our results show that CAPTAIN learns unified multimodal representations by modeling cross-modality dependencies and capturing the diversity of cellular states across complex biological contexts. CAPTAIN generalizes robustly across both fine-tuning and zero-shot settings, excelling in core downstream tasks such as protein imputation and expansion, cell type annotation, and batch harmonization. Beyond improved accuracy in multi-omics integration, CAPTAIN uncovers previously inaccessible mechanisms of protein-driven intercellular dynamics, including immune interaction patterns linked to COVID-19 severity. CAPTAIN establishes a new paradigm for multimodal single-cell modeling, laying the foundation for comprehensive cellular understanding and virtual cell construction. ### Competing Interest Statement The authors have declared no competing interest. We are grateful to members of the Yu laboratory and numerous colleagues for valuable comments and suggestions. This work was supported in part by the grant 2023YFF1204701 from the National Key R\&D Program of China, grant 2024B1515020080 from Guangdong Basic and Applied Basic Research Foundation, grant 32470634 from the National Natural Science Foundation of China, grant KY012023362 from the Talent Research Funding Project of Guangdong Provincial People’s Hospital and grants GZNL2024A01003 and GZNL2023A02002 from the Major Project of Guangzhou National Laboratory. We acknowledge the Data Science Platform of Guangzhou National Laboratory and the Bio-medical Big Data Operating System (Bio-OS) for technical support and for providing access to the computational resources essential to this study.
Cite
CITATION STYLE
Ji, B., Hu, T., Wang, J., Liu, M., Xu, L., Zhang, Q., … Yu, F. (2026). CAPTAIN: a multimodal foundation model pretrained on co-assayed single-cell RNA and protein. Nature Communications. https://doi.org/10.1038/s41467-026-72882-y
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.