NMRExtractor: leveraging large language models to construct an experimental NMR database from open-source scientific publications

11Citations
Citations of this article
21Readers
Mendeley users who have this article in their library.

Abstract

Nuclear magnetic resonance (NMR) spectroscopy is crucial for elucidating molecular structures, but NMR data extraction remains largely manual and time-consuming. We developed NMRExtractor, a locally deployable tool using a fine-tuned large language model, to address this challenge. By processing 5 734 869 open-source scientific publications, we created NMRBank, a dataset containing 225 809 entries with compound IUPAC names, NMR conditions, 1H and 13C NMR chemical shifts, data confidence levels, and reference information. Our analysis reveals that NMRBank's chemical space significantly surpasses existing public NMR datasets. The extraction process is highly scalable, allowing automatic processing of new research papers and continuous updates to NMRBank. This approach not only expands the available open NMR data space but also provides a foundation for AI-based NMR predictions and related chemical research. By automating data extraction and creating a comprehensive, regularly updated NMR database, NMRExtractor and NMRBank address the scarcity of publicly available experimental NMR data, potentially accelerating progress in various fields of chemical research.

Cite

CITATION STYLE

APA

Wang, Q., Zhang, W., Chen, M., Li, X., Xiong, Z., Xiong, J., … Zheng, M. (2025). NMRExtractor: leveraging large language models to construct an experimental NMR database from open-source scientific publications. Chemical Science, 16(25), 11548–11558. https://doi.org/10.1039/d4sc08802f

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free