Abstract
Marine natural products (MNPs) are a diverse group of bioactive compounds with varied chemical structures, but their biological origins are often misannotated due to complex host–microbe symbiosis. Propagated through public databases, such errors hinder biosynthetic studies and AI-driven drug discovery. Here, we develop a structure-based workflow of origin classification and misannotation correction for marine datasets. Using CMNPD and NPAtlas compounds, we integrate a two-step cleaning strategy that detects label inconsistencies and filters structural outliers with a microbial-pretrained graph neural network. The optimized model achieves a balanced accuracy of 85.56% and identifies 3996 compounds whose predicted microbial origins contradict their Animalia labels. These putative symbiotic metabolites cluster within known high-risk taxa, and interpretability analysis reveal biologically coherent structural patterns. This framework provides a scalable quality-control approach for natural product databases and supports more accurate biosynthetic gene cluster (BGC) tracing, host selection, and AI-driven marine natural product discovery.
Author supplied keywords
Cite
CITATION STYLE
Tian, X., Lyu, C., Zhou, Y., Zhang, L., Fan, A., & Liu, Z. (2026). A Structure-Based Deep Learning Framework for Correcting Marine Natural Products’ Misannotations Attributed to Host–Microbe Symbiosis. Marine Drugs, 24(1). https://doi.org/10.3390/md24010020
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.