UniCrop: A Universal, Multi-Source Data Engineering Pipeline for Scalable Crop Yield Prediction

0Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Accurate crop yield prediction increasingly relies on diverse data streams, including satellite observations, meteorological reanalysis, soil composition, and topographic information. However, despite advances in machine learning, many existing approaches remain crop- or region-specific and require substantial bespoke data engineering, limiting scalability and reproducibility. This study introduces UniCrop, a generalisable, configuration-driven data engineering pipeline that standardises the acquisition, harmonisation, and feature construction of multi-source agro-environmental data. Rather than proposing a new predictive model, UniCrop addresses a key bottleneck in agricultural machine learning: the lack of reproducible and scalable data preparation workflows. For any given location, crop type, and temporal window, the pipeline automatically retrieves, harmonises, and engineers over 160 environmental variables from heterogeneous sources (Sentinel-1/2, MODIS, ERA5-Land, NASA POWER, SoilGrids, and SRTM), reducing them to a compact, analysis-ready feature set using a structured feature selection process based on minimum redundancy maximum relevance (mRMR). The effectiveness of the pipeline is demonstrated through a case study, where the generated datasets enable robust baseline modelling across multiple machine-learning algorithms. Using a selected subset of 15 features, four baseline models (LightGBM, Random Forest, Support Vector Regression, and ElasticNet) were evaluated under rigorous cross-validation. LightGBM achieved the best single-model performance (RMSE = 465.1 kg/ha, (Formula presented.) ), while a constrained ensemble provided a marginal improvement (RMSE = 463.2 kg/ha, (Formula presented.) ). SHAP-based analysis further confirms that the selected features capture agronomically meaningful relationships across data modalities. UniCrop contributes a scalable and transparent data engineering pipeline that enables consistent, reproducible, and transferable dataset construction for crop yield prediction. By decoupling data specification from implementation and supporting flexible configuration across crops, regions, and temporal contexts, the framework provides a practical foundation for large-scale agricultural analytics.

Cite

CITATION STYLE

APA

Khidirova, E., & Karakuş, O. (2026). UniCrop: A Universal, Multi-Source Data Engineering Pipeline for Scalable Crop Yield Prediction. Applied Sciences (Switzerland), 16(10). https://doi.org/10.3390/app16104724

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free