Abstract
Motivation Ligands are biomolecules that bind to specific sites on target proteins, often inducing conformational changes important in the protein's function. Knowledge about ligand interactions with proteins are fundamental to understanding biological mechanisms and advancing drug discovery. Traditional protein language models focus on amino acid sequences and 3D structures, overlooking the structural and functional changes induced by protein-ligand interactions. We investigate the value of integrating ligand-protein binding data in several predictive challenges and leverage findings to frame research directions and questions. Results We show how the integration of protein-ligand interaction data in protein representation learning can increase predictive power. We evaluate the methodology across diverse biological tasks, demonstrating consistent improvements over state-of-the-art models. We further demonstrate how the study of the specific boosts in predictive capabilities coming with the introduction of the ligand modality can serve to focus attention and provide insights on biological mechanisms. By leveraging large pretrained protein language models and enriching them with interaction-specific features through a tailored learning process, we capture functional and structural nuances of proteins in their biochemical context. Availability and implementation The full code and data are freely available at https://github.com/kalifadan/ProtLigand (DOI: https://doi.org/10.5281/zenodo.15808053).
Cite
CITATION STYLE
Kalifa, D., Radinsky, K., & Horvitz, E. (2025). Beyond the leaderboard: Leveraging predictive modeling for protein-ligand insights and discovery. Bioinformatics, 41(8). https://doi.org/10.1093/bioinformatics/btaf425
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.