Key points are not available for this paper at this time.
Methods for estimating coffee yield employ expensive and destructive sampling techniques that offer limited flexibility in describing agricultural data variability. The objective of this study was to investigate machine learning (ML) techniques to develop a model that predicts fruit load from nondestructive shoot vegetative growth measurements, integrated with a probabilistic approach for simulating crop yields. Evaluations were conducted on Castillo® Centro variety plants for four consecutive years from sowing to establishment. For ML modeling, data on foliar formation, plagiotropic branch growth, and yield components were collected over three productive years. Three ML techniques were evaluated to estimate fruit load: support vector regression (SVR), artificial neural network (ANN), and random forest (RF). Based on probabilistic distributions from 120 trees, a tree-level yield simulation was conducted, generating a simulated population of 1200 trees. The two most productive branches of each tree were used to parameterize the distributions, incorporating the residual components of the ML models directly into the simulation process. Yield was defined as production per tree in grams (g). The ANN model exhibited the best performance, explaining more than 95% of data variability (R² = 0.98) and the lowest dispersion (RMSE = 3.64 fruits branch⁻¹), with foliar formation contributing 84% to the model structure. The mean difference between simulated and observed yields during the first harvest year did not exceed 300 g per plant. These findings reveal that integrating ML methods with stochastic processes is a robust approach for coffee yield prediction and simulation.
Imbachí et al. (Fri,) studied this question.