Abstract Data from high‐throughput phenotyping (HTP) could be used for phenotype imputation to enhance genomic selection (GS) or gene discovery, but this has not been explored in crop species. Three machine learning models: multiple linear regression (MLR), missForest, and k ‐nearest neighbors, were evaluated for grain yield (GY) phenotype imputation in wheat ( Triticum aestivum L.), using 2414 lines across six environments. Three multispectral vegetation indices (VIs) collected over time from aerial imagery were used as predictors for imputation of simulated missing data ranging from 10% to 70%. Statistical analyses examined the accuracy of imputed GY (IGY) as well as its reliability, heritability, and genetic correlation with GY. Imputation accuracies were highest for MLR, but accuracy differences between methods were small. Accuracies only decreased slightly as percent missing data increased. Genetic correlations between IGYs and observed GY within the environment ranged from −0.01 to 0.51, consistently greater than the corresponding genetic correlations between VIs and GY. Respectively, the reliabilities and heritabilities of IGY were 24% and 45% lower than those of GY, and like those of the VIs. Altogether, this study found that HTP data can be used to impute GY phenotypes suitable for use in further analyses; however, imputed and observed GY should be modeled as separate traits. Further research is needed to improve the heritability of IGY and to evaluate the utility of IGY in GS models.
Gevartosky et al. (Tue,) studied this question.