← Papers

Unverified paper record

Transcriptome-based prediction for polygenic traits in rice using different gene subsets.

BMC genomics · 1 Oct 2024 · 10.1186/s12864-024-10803-3

Abstract

Background Transcriptome-based prediction of complex phenotypes is a relatively new statistical method that links genetic variation to phenotypic variation. The selection of large-effect genes based on a priori biological knowledge is beneficial for predicting oligogenic traits; however, such a simple gene selection method is not applicable to polygenic traits because causal genes or large-effect loci are often unknown. Here, we used several gene-level features and tested whether it was possible to select a gene subset that resulted in better predictive ability than using all genes for predicting a polygenic trait. Results Using the phenotypic values of shoot and root traits and transcript abundances in leaves and roots of 57 rice accessions, we evaluated the predictive abilities of the transcriptome-based prediction models. Leaf transcripts predicted shoot phenotypes, such as plant height, more accurately than root transcripts, whereas root transcripts predicted root phenotypes, such as crown root length, more accurately than leaf transcripts. Furthermore, we used the following three features to train the prediction model: (1) tissue specificity of the transcripts, (2) ontology annotations, and (3) co-expression modules for selecting gene subsets. Although models trained by a gene subset often resulted in lower predictive abilities than the model trained by all genes, some gene subsets showed improved predictive ability. For example, using genes expressed in roots but not in leaves, the predictive ability for crown root diameter was improved by more than 10% (R 2 = 0.59 when using all genes; R 2 = 0.66, using 1,554 root-specifically expressed genes). Similarly, genes annotated as "gibberellic acid sensitivity" showed higher predictive ability than using all genes for root dry weight. Conclusions Our results highlight both the possibility and difficulty of selecting an appropriate gene subset to predict polygenic traits from transcript abundance, given the current biological knowledge and information. Further integration of multiple sources of information, as well as improvements in gene characterization, may enable the selection of an optimal gene set for the prediction of polygenic phenotypes.

Plant phenotyping relevance

遺伝子発現データから植物の複合形質を予測する統計的手法を開発・評価しており、形質予測モデルの性能比較が研究の中心である。

abstractTranscriptome-based prediction of complex phenotypes is a relatively new statistical method that links genetic variation to phenotypic variation.
abstractwe used several gene-level features and tested whether it was possible to select a gene subset that resulted in better predictive ability than using all genes for predicting a polygenic trait.
abstractwe evaluated the predictive abilities of the transcriptome-based prediction models.

Code and data availability

The authors state that all analysis code for the transcriptome-based prediction study is publicly available on Figshare, which is a paper-specific, publicly actionable code asset. Phenotype data are only in a prior study's supplementary file and transcriptome data in GEO (GSE162313), which are cited prior deposits, not

Codepublic

w sequence data were deposited in the DNA Data Bank of Japan Sequence Read Archive in a previous study [27]. Transcriptome data are available from the Gene Expression Omnibus ( GSE162313 ) and phenotype data are available in the supplementary file of a previous study [28]. All codes for the data analysis are shown in Figshare ( https://doi.org/10.6084/m9.figshare.26067532.v1 ). Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing interests The authors declare no competing interests. Abbreviations WRC

Open resource ↗Figshare · 10.6084/m9.figshare.26067532.v1 · lines:370-396

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.