← Papers

Unverified paper record

From genotype to phenotype in Arabidopsis thaliana: in-silico genome interpretation predicts 288 phenotypes from sequencing data.

Nucleic acids research · 1 Feb 2022 · 10.1093/nar/gkab1099

Abstract

In many cases, the unprecedented availability of data provided by high-throughput sequencing has shifted the bottleneck from a data availability issue to a data interpretation issue, thus delaying the promised breakthroughs in genetics and precision medicine, for what concerns Human genetics, and phenotype prediction to improve plant adaptation to climate change and resistance to bioagressors, for what concerns plant sciences. In this paper, we propose a novel Genome Interpretation paradigm, which aims at directly modeling the genotype-to-phenotype relationship, and we focus on A. thaliana since it is the best studied model organism in plant genetics. Our model, called Galiana, is the first end-to-end Neural Network (NN) approach following the genomes in/phenotypes out paradigm and it is trained to predict 288 real-valued Arabidopsis thaliana phenotypes from Whole Genome sequencing data. We show that 75 of these phenotypes are predicted with a Pearson correlation ≥0.4, and are mostly related to flowering traits. We show that our end-to-end NN approach achieves better performances and larger phenotype coverage than models predicting single phenotypes from the GWAS-derived known associated genes. Galiana is also fully interpretable, thanks to the Saliency Maps gradient-based approaches. We followed this interpretation approach to identify 36 novel genes that are likely to be associated with flowering traits, finding evidence for 6 of them in the existing literature.

Plant phenotyping relevance

植物の遺伝子型から288形質を予測するエンドツーエンドのニューラルネットワーク手法を開発・評価しており、表現型推定手法が研究の中心である。

abstractOur model, called Galiana, is the first end-to-end Neural Network (NN) approach following the genomes in/phenotypes out paradigm and it is trained to predict 288 real-valued Arabidopsis thaliana phenotypes from Whole Genome sequencing data.
abstractWe show that 75 of these phenotypes are predicted with a Pearson correlation ≥0.4

Code and data availability

The paper's authors explicitly state that the Galiana model code is freely available from their public Bitbucket repository, which is a paper-specific computational analysis asset. The phenotype data come from third-party databases (1001 Genomes, AraPheno) and are not paper-specific deposits; supplementary tables are a

Codepublic

We implemented the model using pytorch ( 26 ). The code is freely available from our git repository https://bitbucket.org/eddiewrc/galiana/src/master/ .

Open resource ↗bitbucket.org/eddiewrc/galiana · eddiewrc/galiana · lines:44-55

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.