← Papers

Unverified paper record

A comparison of classical and machine learning-based phenotype prediction methods on simulated data and three plant species

Frontiers in Plant Science · 4 Nov 2022 · 10.3389/fpls.2022.932512

Abstract

Genomic selection is an integral tool for breeders to accurately select plants directly from genotype data leading to faster and more resource-efficient breeding programs. Several prediction methods have been established in the last few years. These range from classical linear mixed models to complex non-linear machine learning approaches, such as Support Vector Regression, and modern deep learning-based architectures. Many of these methods have been extensively evaluated on different crop species with varying outcomes. In this work, our aim is to systematically compare 12 different phenotype prediction models, including basic genomic selection methods to more advanced deep learning-based techniques. More importantly, we assess the performance of these models on simulated phenotype data as well as on real-world data from Arabidopsis thaliana and two breeding datasets from soy and corn. The synthetic phenotypic data allow us to analyze all prediction models and especially the selected markers under controlled and predefined settings. We show that Bayes B and linear regression models with sparsity constraints perform best under different simulation settings with respect to explained variance. Further, we can confirm results from other studies that there is no superiority of more complex neural network-based architectures for phenotype prediction compared to well-established methods. However, on real-world data, for which several prediction models yield comparable results with slight advantages for Elastic Net, this picture is less clear, suggesting that there is a lot of room for future research.

Plant phenotyping relevance

複数の表現型予測モデルを植物種の実データとシミュレーションで系統比較・評価しており、計算による植物形質推定が研究の中心である。

abstractour aim is to systematically compare 12 different phenotype prediction models
abstractwe assess the performance of these models on simulated phenotype data as well as on real-world data from Arabidopsis thaliana and two breeding datasets from soy and corn

Code and data availability

The paper's authors publicly release their analysis code (easyPheno framework and the phenotype_prediction repository containing simulated phenotypes, hyperparameter optimization results, GWAS results, and figure-generation code), plus the Arabidopsis SNP matrix (figshare) and four AraPheno phenotype datasets used in a

Codepublic

All simulated phenotypes, detailed results of the whole hyperparameter optimization, precomputed permutation-based GWAS results, and the code for conducting the simulations and generating all figures can be freely downloaded from our GitHub repository: https://github.com/grimmlab/phenotype_prediction .

Open resource ↗grimmlab/phenotype_prediction · lines:619-631
Datasetpublic

The fully imputed SNP matrix data for Arabidopsis thaliana is publicly available and can be downloaded from https://doi.org/10.6084/m9.figshare.11346893.v1 .

Open resource ↗10.6084/m9.figshare.11346893.v1 · lines:619-631

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.