inversely proportionate to expected model error was 131 optimal for these data. 132 133 Materials and Methods 134 Data Preparation 135 Maize yield, environmental, and management data came from the Genomes to Fields 136 (G2F) initiative’s data releases for 2014-2019 (McFarland et al. 2020). These data are publicly 137 available (https://www.genomes2fields.org/resources/) and provide weather and soil data for 138 fields in the continental United States, management information, and genomic data in addition to 139 phenotypic measurements. Weather data was supplemented with data from Daymet (Thornton 140 et al. 2020). Genomic, environmental, and management data was quality controlled using 141 cu
Open resource ↗Genomes to Fields · pdf-raw-page:6 lines:1-81Unverified paper record
Ensemble of BLUP, Machine Learning, and Deep Learning Models Predict Maize Yield Better Than Each Model Alone
6 Apr 2023 · 10.21203/rs.3.rs-2757632/v1
Abstract
Abstract Predicting phenotypes accurately from genomic, environment, and management factors is key to accelerating the development of novel cultivars with desirable traits. Inclusion of management and environmental factors enables in silico studies to predict the effect of specific management interventions or future climates. Despite the value such models would confer, much work remains to improve the accuracy of phenotypic predictions. Rather than advocate for a single specific modeling strategy, here we demonstrate within large multi-environment and multi-genotype maize trials that combining predictions from disparate models using simple ensemble approaches most often results in better accuracy than using any one of the models on their own. We investigated various ensemble combinations of different model types, model numbers, and model weighting schemes to determine the accuracy of each. We find that ensembling generally improves performance even when combining only two models. The number and type of models included alter accuracy with improvements diminishing as the number of models included increases. Using a genetic algorithm to optimize ensemble composition reveals that, when weighted by the inverse of each model’s expected error, using combinations of best linear unbiased predictors, linear fixed effects models, deep learning models, and select machine learning models perform best on our datasets.
Plant phenotyping relevance
トウモロコシ収量という植物形質を対象に、BLUP・機械学習・深層学習のアンサンブル予測を比較・検証し、予測精度を評価することが中心であるため。
abstractPredicting phenotypes accurately from genomic, environment, and management factors is key to accelerating the development of novel cultivars with desirable traits.
abstractwe demonstrate within large multi-environment and multi-genotype maize trials that combining predictions from disparate models using simple ensemble approaches most often results in better accuracy than using any one of the models on their own.
Code and data availability
The paper's maize yield, weather, soil, management, and genomic data come from the Genomes to Fields initiative's public releases (2014-2019), accessible via the G2F resources page. The authors also deposit cleaned data (Zenodo 10.5281/zenodo.6916775) and analysis scripts/model predictions (Zenodo 10.5281/zenodo.769738
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.