Abstract
Nutrients play a critical role in oceanic primary productivity and the biological pump. However, compared to hydrographic parameters such as temperature and salinity, nutrient observations are limited due to their labor-intensive and costly measurements. Thus, nutrient observations are several orders of magnitude sparser than hydrographic observations. In this study, we first established a rigorous data quality control procedure to clean the hydrographic and nutrient (including NO3-, NO2-, DIP, and Si(OH)4) observations collected from World Ocean Database (WOD) and CLIVAR and Carbon Hydrographic Data Office (CCHDO) in the North Pacific. Subsequently, the cleaned and high-quality CCHDO dataset was used to train three machine learning models - Random Forest, Light Gradient Boosting Machine (LightGBM), and Gaussian Process Regression - to establish relationships between nutrient concentrations and key variables, including space coordinates (longitude, latitude, and depth), time variables (year and month), and water mass properties (indexed by potential temperature and salinity). Validation shows that the reconstruction closely matches the observations, with Root Mean Squared Errors (RMSEs) of <1.41, <0.071, <0.089 and <3.07 μmol kg-1 for NO3-, NO2-, DIP, and Si(OH)4, respectively. The validated models were then applied to reconstruct nutrient concentrations from the hydrographic observations in WOD, most of which lacked direct nutrient measurements. This resulted in ∼ 473 million reconstructed nutrient data points across 1.92 million stations for each nutrient, spanning from 1895 to 2024, representing a 2127- to 2393-fold increase compared to the original nutrient observations in the North Pacific (197 539 to 222 234). This new dataset will be valuable for studying nutrient transport and budgets, spinning up and validating ocean biogeochemical models, assessing long-term nutrients and their stoichiometric changes driven by anthropogenic forcing and climate change. The dataset generated in this study is openly available via Zenodo (10.5281/zenodo.17451417) (Du et al., 2025).
Cite
CITATION STYLE
Du, C., Zheng, N., Kao, S. J., Dai, M., Cao, Z., Shi, D., … Li, X. (2026). A historical nutrient dataset (1895-2024) for the North Pacific: reconstructed from machine learning and hydrographic observations. Earth System Science Data, 18(4), 2951–2969. https://doi.org/10.5194/essd-18-2951-2026
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.