← Papers

Unverified paper record

Improved genomic prediction performance with ensembles of diverse models.

G3 (Bethesda, Md.) · 1 May 2025 · 10.1093/g3journal/jkaf048

Abstract

The improvement of selection accuracy of genomic prediction is a key factor in accelerating genetic gain for crop breeding. Traditionally, efforts have focused on developing superior individual genomic prediction models. However, this approach has limitations due to the absence of a consistently "best" individual genomic prediction model, as suggested by the No Free Lunch Theorem. The No Free Lunch Theorem states that the performance of an individual prediction model is expected to be equivalent to the others when averaged across all prediction scenarios. To address this, we explored an alternative method: combining multiple genomic prediction models into an ensemble. The investigation of ensembles of prediction models is motivated by the Diversity Prediction Theorem, which indicates the prediction error of the many-model ensemble should be less than the average error of the individual models due to the diversity of predictions among the individual models. To investigate the implications of the No Free Lunch and Diversity Prediction Theorems, we developed a naïve ensemble-average model, which equally weights the predicted phenotypes of individual models. We evaluated this model using 2 traits influencing crop yield-days to anthesis and tiller number per plant-in the teosinte nested association mapping dataset. The results show that the ensemble approach increased prediction accuracies and reduced prediction errors over individual genomic prediction models. The advantage of the ensemble was derived from the diverse predictions among the individual models, suggesting the ensemble captures a more comprehensive view of the genomic architecture of these complex traits. These results are in accordance with the expectations of the Diversity Prediction Theorem and suggest that ensemble approaches can enhance genomic prediction performance and accelerate genetic gain in crop breeding programs.

Plant phenotyping relevance

作物の表現型形質を予測するアンサンブル計算法の開発・評価が中心であり、単なる育種実験ではないため、計算的な表現型推定手法として収載する。

abstractwe developed a naïve ensemble-average model, which equally weights the predicted phenotypes of individual models.
abstractWe evaluated this model using 2 traits influencing crop yield-days to anthesis and tiller number per plant

Code and data availability

The paper's Data availability statement explicitly deposits the authors' analysis code on GitHub and the data and code on Zenodo. The phenotype/genotype data itself (TeoNAM) was collected by Chen et al. (2019) and is cited as publicly available, but the Zenodo record contains the data and code used in this study.

Codepublic

The code generated for this experiment is shared at https://github.com/ShunichiroT/ensemble .

Open resource ↗https://github.com/ShunichiroT/ensemble · lines:323-356
Datasetpublic

The data and code used in this study were also uploaded at https://zenodo.org/records/14776591 .

Open resource ↗https://zenodo.org/records/14776591 · lines:323-356

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.