Genomic selection (GS) estimates the GEBV from genome-wide markers to reduce generation intervals and optimize germplasm selection, which is particularly advantageous for high-cost or late-expressed traits. While models like GBLUP are popular, they assume a polygenic architecture. In contrast, the Bayesian alphabet and machine learning (ML) can accommodate other types of genetic architectures. Given that no single model is universally optimal, stacking ensembles, which train a meta-model using predictions from diverse base learners, emerge as a compelling solution. However, the application of stacking in GS often overlooks non-additive effects. This study evaluated different stacking configurations for genomic prediction across 10 simulated traits, covering additive, dominance, and epistatic genetic architectures. A 5-fold cross-validation scheme was used to assess predictive ability and other evaluation metrics. The stacking approach demonstrated superior predictive ability in all scenarios. Gains were especially pronounced in complex architectures (100 QTLs, h2 = 0.3), reaching an 83% increment over the best individual model (BayesA with dominance), and also in oligogenic scenarios with epistasis (10 QTLs, h2 = 0.6), with a 27.59% gain. The success of stacking was attributed to two key strategies: base learner selection and the use of robust meta-learners (such as principal component or penalized regression) that effectively handled multicollinearity.
Celeri et al. (Tue,) studied this question.