← Papers

Unverified paper record

Predicting Phenotypic Traits Using a Massive RNA-seq Dataset

7 Dec 2023 · 10.1101/2023.12.05.570195

Abstract

Transcriptomic data can be used to predict environmentally impacted phenotypic traits. This type of prediction is particularly useful for monitoring difficult-to-measure phenotypic traits and has become increasingly popular for monitoring high-value agricultural crops and in precision medicine. Despite this increase in popularity, little research has been done on how many samples are required for these models to be accurate, and which normalization should be used. Here we create a massive RNA-seq dataset from publicly available Arabidopsis thaliana data with corresponding measurements for age and tissue type. We use this dataset to determine how many samples are required for accurate model prediction and which normalization method is required. We find that Median Ratios Normalization significantly increases performance when predicting age. We also find that in the case of our dataset, only a few hundred samples are required to predict tissue types, and only a few thousand samples are necessary to accurately predict age. Researchers should consider these results when choosing the number of samples in a transcriptomic experiment and during data-processing. Author Summary Large datasets have become ubiquitous in both research and industry, with thousands and sometimes millions of samples being collected for a single project. In biology a prominent new technology is RNA-seq, which can be used to measure the expression level of thousands of genes for a single sample. These measurements are used for a variety of downstream applications, including predicting phenotypic traits (i.e. height, disease, etc.). A number of experiments have attempted to use RNA-seq data to make phenotype predictions with varying success. This is partially due to the small sample size of their experiments. RNA-seq datasets are currently relatively small--only a dozen to a few hundred samples--due to the cost per sample. This is expected to change as the cost of sequencing decreases. In this paper we create a massive conglomerate RNA-seq dataset from publicly available Arabidopsis thaliana RNA-seq data. We use this dataset to determine how many samples are required to accurately predict plant age and tissue type using machine learning models. We also explore the best way to normalize large datasets. Our results show the potential of massive RNA-seq datasets, and can be used to inform experimental design for phenotype prediction.

Plant phenotyping relevance

RNA-seqデータから植物の年齢・組織型を予測する機械学習について、必要サンプル数と正規化法を大規模Arabidopsisデータで評価しており、表現型推定手法の検証が中心である。

abstractWe use this dataset to determine how many samples are required for accurate model prediction and which normalization method is required.
abstractWe use this dataset to determine how many samples are required to accurately predict plant age and tissue type using machine learning models.
abstractOur results show the potential of massive RNA-seq datasets, and can be used to inform experimental design for phenotype prediction.

Code and data availability

The paper's normalized gene expression matrices, curated phenotype annotation datasets, and intermediary files are publicly deposited on Zenodo, and all analysis code is publicly available on GitLab. These directly reproduce the paper's plant-phenotyping measurements (Arabidopsis age/tissue annotations) and modeling/ML

Datasetpublic

(NoNo). TMM normalization [24,29] and MRN normalization [25] were performed using the Python “conorm” package 1.2.0 [30]. TPM and NoNo normalization values were an output of Kallisto [27]. How these normalizations impacted sample count is visualized as S2 Figure. We have made these GEMs publicly available on Zenodo at the link https://zenodo.org/records/10183151 Sample Phenotype Annotations Pre-Processing Sample phenotype annotations were retrieved from the NCBI BioProject database [16,17] using BioSampleParser which was slightly modified to check for successful data retrieval [31]. Phenotype annotations were retrieved for 48696 NCBI BioSamples, representing data from 2643 BioProjects.

Open resource ↗Zenodo · pdf-raw-page:9 lines:1-55
Datasetpublic

Data Availability Statement All normalized gene expression datasets, phenotype datasets, and intermediary files created for this research are publically available on Zenodo at link https://zenodo.org/doi/10.5281/zenodo.10183150 All code written in support of this publication is publicly available on GitLab at link https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics Funding This work was supported by the Washington Tree Fruit Research Commission (WTFRC) project #AP-22-101 and USDA ARS internal appropriation funds. References 1. Bostanci

Open resource ↗Zenodo · 10.5281/zenodo.10183150 · pdf-raw-page:43 lines:1-51
Codepublic

Data Availability Statement All normalized gene expression datasets, phenotype datasets, and intermediary files created for this research are publically available on Zenodo at link https://zenodo.org/doi/10.5281/zenodo.10183150 All code written in support of this publication is publicly available on GitLab at link https://gitlab.com/ficklinlab-public/modeling-with-transcriptomics Funding This work was supported by the Washington Tree Fruit Research Commission (WTFRC) project #AP-22-101 and USDA ARS internal appropriation funds. References 1. Bostanci E, Kocak E, Unal M, Guzel MS, Acici K, Asuroglu T. Machine Learning Analysis of RNA-seq Data for Diagnostic and Prognostic Prediction of Colon

Open resource ↗GitLab · pdf-raw-page:43 lines:1-51

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.