← Papers

Unverified paper record

Automated Methods Enable Direct Computation on Phenotypic Descriptions for Novel Candidate Gene Prediction.

Frontiers in Plant Science · 9 Jan 2020 · 10.3389/fpls.2019.01629

Abstract

Natural language descriptions of plant phenotypes are a rich source of information for genetics and genomics research. We computationally translated descriptions of plant phenotypes into structured representations that can be analyzed to identify biologically meaningful associations. These representations include the entity-quality (EQ) formalism, which uses terms from biological ontologies to represent phenotypes in a standardized, semantically rich format, as well as numerical vector representations generated using natural language processing (NLP) methods (such as the bag-of-words approach and document embedding). We compared resulting phenotype similarity measures to those derived from manually curated data to determine the performance of each method. Computationally derived EQ and vector representations were comparably successful in recapitulating biological truth to representations created through manual EQ statement curation. Moreover, NLP methods for generating vector representations of phenotypes are scalable to large quantities of text because they require no human input. These results indicate that it is now possible to computationally and automatically produce and populate large-scale information resources that enable researchers to query phenotypic descriptions directly.

Plant phenotyping relevance

植物表現型記述をNLPで構造化・ベクトル化し、手動キュレーションとの性能比較で検証する計算手法が中心である。

abstractWe computationally translated descriptions of plant phenotypes into structured representations that can be analyzed to identify biologically meaningful associations.
abstractWe compared resulting phenotype similarity measures to those derived from manually curated data to determine the performance of each method.

Code and data availability

The paper's data availability statement explicitly deposits the authors' analysis code on GitHub (irbraun/phenologs) and all files needed to reproduce the results (including the phenotype/EQ datasets used) on Zenodo (doi 10.5281/zenodo.3255020). These are paper-specific, public, and actionable assets for the phenotype-

Codepublic

The code used to produce the results of this work is available at github.com/irbraun/phenologs . Files necessary to reproduce the discussed results, datasets used to generate figures presented in this work, and other supplemental files are available at doi.org/10.5281/zenodo.3255020 .

Open resource ↗irbraun/phenologs · 10.5281/zenodo.3255020 · lines:792-814
Datasetpublic

Files necessary to reproduce the discussed results, datasets used to generate figures presented in this work, and other supplemental files are available at doi.org/10.5281/zenodo.3255020 . This data repository also includes versions of the previously described datasets available as supplemental data of Oellrich, Walls et al. (2015) and Lloyd and Meinke (2012) , for the purpose of making this study reproducible without any additional external files.

Open resource ↗10.5281/zenodo.3255020 · lines:792-814

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.