Unverified paper record
Genomic and Phenomic Prediction for Soybean Seed Yield, Protein, and Oil
bioRxiv · 2 Nov 2024 · 10.1101/2024.11.01.621550
Abstract
Developments in genomics and phenomics have provided valuable tools for use in cultivar development. Genomic prediction (GP) has been used in commercial soybean [Glycine max L. (Merr.)] breeding programs to predict grain yield and seed composition traits. Phenomic prediction (PP) is a rapidly developing field that holds the potential to be used for the selection of genotypes early in the growing season. The objectives of this study were to compare the use and performance of GP and PP for predicting soybean seed yield, protein content, and oil content. We additionally conducted Genome Wide Association Studies (GWAS) to identify significant SNPs associated with the traits of interest. These SNPs were also used to train the GP models. The GWAS panel of 292 diverse accessions was grown in six environments in replicated trials. Spectral data were collected at three timepoints during the growing season. A GBLUP model was trained on 268 accessions, while three separate machine learning (ML) models were trained on vegetation indices (VIs) and canopy traits. We observed that for PP, Random Forest (RF) algorithm had the highest rank correlation between the predicted and the actual phenotype rank. PP had a higher correlation coefficient than GP for seed yield, while GP had higher correlation coefficients for seed protein and oil contents. VIs with high feature importance were used as covariates in a new GBLUP model, and a new RF model was trained with the inclusion of selected SNPs from the GWAS results. These models did not outperform the original GP and PP models. These results show the capability of using ML for in-season predictions for specific traits in soybean breeding and provide insights on PP and GP inclusions in breeding programs.
Plant phenotyping relevance
スペクトルデータ、植生指数、キャノピー形質、機械学習を用いたフェノミック予測を中心に、収量・種子成分の予測性能をゲノム予測と比較しているため、植物表現型取得・推定手法の実質的な評価に該当する。
abstractThe objectives of this study were to compare the use and performance of GP and PP for predicting soybean seed yield, protein content, and oil content.
abstractSpectral data were collected at three timepoints during the growing season.
abstractWe observed that for PP, Random Forest (RF) algorithm had the highest rank correlation between the predicted and the actual phenotype rank.
abstractThese results show the capability of using ML for in-season predictions for specific traits in soybean breeding
Code and data availability
The supplied blocks describe soybean phenomic/genomic prediction methods and results but contain no data availability statement, deposit accession, or author code repository URL. The only public resource mentioned (SoySNP50K genotype data) is a cited prior dataset, not a paper-specific asset. No qualifying public asset
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.