Accurate estimation of agricultural output is vital for food security and resource management, especially in agrarian economies like India. Punjab and Haryana, key rice-producing states, face challenges with traditional yield assessment methods, which are often slow, labor-intensive, and limited in scope. This study presents a machine learning-based framework to predict paddy yield using over four decades (1981-2023) of historical agro-meteorological and crop production data. Key climatic variables such as rainfall, temperature, solar radiation, and soil moisture were combined with past yield and acreage statistics to train models, including Random Forest, LSTM, and Monte Carlo dropout. The models were evaluated against 2023 yield data to test predictive accuracy and generalizability. Among the evaluated models, the LSTM network achieved the highest deterministic accuracy, with R² of 0.91 for both states and MAPE values of 3.2% for Haryana and 5.28% for Punjab, representing a 74-86% improvement over the Linear Regression baseline. The LSTM model combined with Monte Carlo dropout achieved R² values of 0.96 in Haryana and 0.88 in Punjab, with corresponding MAPE values of 2.1% and 7.43%, respectively, and generated 95% prediction intervals of ±8-12% for Haryana and ±12-18% for Punjab. Although the probabilistic framework did not uniformly improve point-estimate accuracy across both states, it provided uncertainty bounds that enhance the utility of yield forecasts for risk-aware agricultural planning and decision-making. The results suggest that variables, such as rainfall, soil moisture, and thermal stress, play an important role in explaining yield variability. These findings demonstrate that long-term agro-meteorological records can support accurate, scalable, and uncertainty-aware yield forecasting, providing a robust data-driven alternative to conventional field-based assessment methods for agricultural planning and policymaking.
Khare et al. (Sun,) studied this question.