Accurate crop yield prediction increasingly relies on diverse data streams, including satellite observations, meteorological reanalysis, soil composition, and topographic information. However, despite advances in machine learning, many existing approaches remain crop- or region-specific and require substantial bespoke data engineering, limiting scalability and reproducibility. This study introduces UniCrop, a generalisable, configuration-driven data engineering pipeline that standardises the acquisition, harmonisation, and feature construction of multi-source agro-environmental data. Rather than proposing a new predictive model, UniCrop addresses a key bottleneck in agricultural machine learning: the lack of reproducible and scalable data preparation workflows. For any given location, crop type, and temporal window, the pipeline automatically retrieves, harmonises, and engineers over 160 environmental variables from heterogeneous sources (Sentinel-1/2, MODIS, ERA5-Land, NASA POWER, SoilGrids, and SRTM), reducing them to a compact, analysis-ready feature set using a structured feature selection process based on minimum redundancy maximum relevance (mRMR). The effectiveness of the pipeline is demonstrated through a case study, where the generated datasets enable robust baseline modelling across multiple machine-learning algorithms. Using a selected subset of 15 features, four baseline models (LightGBM, Random Forest, Support Vector Regression, and ElasticNet) were evaluated under rigorous cross-validation. LightGBM achieved the best single-model performance (RMSE = 465.1 kg/ha, R2=0.6576), while a constrained ensemble provided a marginal improvement (RMSE = 463.2 kg/ha, R2=0.6604). SHAP-based analysis further confirms that the selected features capture agronomically meaningful relationships across data modalities. UniCrop contributes a scalable and transparent data engineering pipeline that enables consistent, reproducible, and transferable dataset construction for crop yield prediction. By decoupling data specification from implementation and supporting flexible configuration across crops, regions, and temporal contexts, the framework provides a practical foundation for large-scale agricultural analytics.
Khidirova et al. (Sun,) studied this question.