← Papers

Unverified paper record

Enhancing genome-wide populus trait prediction through deep convolutional neural networks.

The Plant journal : for cell and molecular biology · 13 May 2024 · 10.1111/tpj.16790

Abstract

As a promising model, genome-based plant breeding has greatly promoted the improvement of agronomic traits. Traditional methods typically adopt linear regression models with clear assumptions, neither obtaining the linkage between phenotype and genotype nor providing good ideas for modification. Nonlinear models are well characterized in capturing complex nonadditive effects, filling this gap under traditional methods. Taking populus as the research object, this paper constructs a deep learning method, DCNGP, which can effectively predict the traits including 65 phenotypes. The method was trained on three datasets, and compared with other four classic models-Bayesian ridge regression (BRR), Elastic Net, support vector regression, and dualCNN. The results show that DCNGP has five typical advantages in performance: strong prediction ability on multiple experimental datasets; the incorporation of batch normalization layers and Early-Stopping technology enhancing the generalization capabilities and prediction stability on test data; learning potent features from the data and thus circumventing the tedious steps of manual production; the introduction of a Gaussian Noise layer enhancing predictive capabilities in the case of inherent uncertainties or perturbations; fewer hyperparameters aiding to reduce tuning time across datasets and improve auto-search efficiency. In this way, DCNGP shows powerful predictive ability from genotype to phenotype, which provide an important theoretical reference for building more robust populus breeding programs.

Plant phenotyping relevance

ポプラの65形質を遺伝子型から予測する深層学習手法DCNGPを開発し、複数データセットおよび既存モデルと比較評価しており、植物形質推定手法が研究の中心である。

abstractthis paper constructs a deep learning method, DCNGP, which can effectively predict the traits including 65 phenotypes.
abstractThe method was trained on three datasets, and compared with other four classic models-Bayesian ridge regression (BRR), Elastic Net, support vector regression, and dualCNN.

Code and data availability

The paper's authors publicly released the DCNGP analysis scripts on GitHub; the SRA deposits contain only raw sequencing reads (molecular omics), not phenotype datasets, so they are excluded.

Codepublic

ILABILITY STATEMENT All raw sequencing reads have been deposited in NCBI’s Sequence Read Archive (SRA) under accession number PRJNA297202 and PRJNA510671 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA297202/, https://www.ncbi.nlm.nih.gov/bioproject/PRJNA510671/). The DCNGP scripts are available in the release package on GitHub (https://github.com/xiangweidai/DCNGP).SUPPORTING INFORMATION Additional Supporting Information may be found in the online ver- sion of this article. Figure S1. All 65 predicted phenotypes and their abbreviations in this work. Table S1. Predictions of 30 phenotypes from 94 samples by four DL models in ablation experiment. Table S2. Architecture details of the deep-

Open resource ↗xiangweidai/DCNGP · pdf-raw-page:10 lines:1-94

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.