← Papers

Unverified paper record

Predicting adult phenotypes from seedling transcriptional data using deep learning: a case study in chrysanthemum

Genomics Communications · 1 Jan 2026 · 10.48130/gcomm-0026-0011

Abstract

Genotype-to-phenotype prediction remains a fundamental challenge in current genetic research. In recent years, it has become possible to construct different predictive models based on genomic data. However, in many horticultural crops, it is difficult to accurately verify genomic variations because of the complexity of their genome, making the application of these genome-based methods challenging. Gene expression reflects both genetic regulatory mechanisms and environmental stimuli, offering potential for predicting phenotypes in plants with complex genomes. Thus, in this paper, we tested the possibility for predicting adult plant phenotypes using the gene expression data from seedlings. By applying the transcriptional-based deep learning methods on cut chrysanthemums (Chrysanthemum spp.), which exhibits a complex genetic background characterized by high repetitiveness, heterozygosity, and genome size and is recognized as a segmental allopolyploid, we found that the method is robust and accurate for predicting continuous variables such as leaf vase life, as well as categorical variables such as flower types on the basis of gene expression data. Moreover, the power and performance of transcriptional-based deep learning methods for prediction was validated in rice (Oryza sativa). Our research shows the good performance of phenotype prediction based on gene expression, with potential applications in future gene chip-based breeding practices.

Plant phenotyping relevance

遺伝子発現データから成体の植物形質を予測する深層学習手法を開発・検証しており、形質予測が研究の中心である。

titlePredicting adult phenotypes from seedling transcriptional data using deep learning: a case study in chrysanthemum
abstractMoreover, the power and performance of transcriptional-based deep learning methods for prediction was validated in rice (Oryza sativa).

Code and data availability

The paper deposits its authors' analysis code publicly on GitHub and its raw RNA-seq data (used for the seedling-transcriptome phenotype prediction) in the Genome Sequence Archive with accession CRA022074. Both are paper-specific, public, and actionable.

Codepublic

n for multiclass classification. For compiling each model, the RMSprop optimization algorithm was used with a default initial learning rate of 0.001, and categorical cross-entropy was selected as the loss func- tion. The model was trained for 100 epochs with a default batch size of 32. The source codes are publicly available at https://github.com/lkwwang-ui/Deep-model-for-predicting-adult-traits-using-seedling-data-study.git We used Weka 3.9.7 data mining software[23] and performed machine learning analysis as described in our previously published paper[24]. In brief, all 101 samples were used for training and testing with 10-fold cross-validation, and the 20 samples from BGZ were used for m

Open resource ↗https://github.com/lkwwang-ui/Deep-model-for-predicting-adult-traits-using-seedling-data-study.git · pdf-raw-page:3 lines:1-80

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.