The dataset used in this work is available at https://git.io/JTutQ.
Open resource ↗pdf-page:1 lines:1-60Unverified paper record
The Case for Retaining Natural Language Descriptions of Phenotypes in Plant Databases and a Web Application as Proof of Concept
bioRxiv (Cold Spring Harbor Laboratory) · 6 Feb 2021 · 10.1101/2021.02.04.429796
Abstract
ABSTRACT Similarities in phenotypic descriptions can be indicative of shared genetics, metabolism, and stress responses, to name a few. Finding and measuring similarity across descriptions of phenotype is not straightforward, with previous successes in computation requiring a great deal of expert data curation. Natural language processing of free text descriptions of phenotype is often less resource intensive than applying expert curation. It is therefore critical to understand the performance of natural language processing techniques for organizing and analyzing biological datasets and for enabling biological discovery. For predicting similar phenotypes, a wide variety of approaches from the natural language processing domain perform as well as curation-based methods. These computational approaches also show promise both for helping curators organize and work with large datasets and for enabling researchers to explore relationships among available phenotype descriptions. Here we generate networks of phenotype similarity and share a web application for querying a dataset of associated plant genes using these text mining approaches. Example situations and species for which application of these techniques is most useful are discussed. Database URLs The database and analytical tool called QuOATS are available at https://quoats.dill-picl.org/ . Code for the web application is available at https://git.io/Jtv9J . Datasets are available for direct access via https://zenodo.org/record/7947342#.ZGwAKOzMK3I . The code for the analyses performed for the publication is available at https://github.com/Dill-PICL/Plant-data and https://github.com/Dill-PICL/NLP-Plant-Phenotypes .
Plant phenotyping relevance
植物表現型の自然言語記述をNLPで類似性解析し、遺伝子データセット探索用のWebアプリケーションを開発・提供しており、表現型データの計算的整理・解析手法が中心です。
abstractNatural language processing of free text descriptions of phenotype is often less resource intensive than applying expert curation.
abstractHere we generate networks of phenotype similarity and share a web application for querying a dataset of associated plant genes using these text mining approaches.
Code and data availability
The paper's phenotype description dataset and analysis code are explicitly deposited with public git.io URLs, and the QuOATS web application is publicly hosted. All are paper-specific and actionable.
The code for the analysis performed here is available at https://git.io/JTutN and https://git.io/JTuqv.
Open resource ↗pdf-page:1 lines:1-60The code for the analysis performed here is available at https://git.io/JTutN and https://git.io/JTuqv.
Open resource ↗pdf-page:1 lines:1-60The code for the web application discussed here is available at https://git.io/Jtv9J, and the application itself is available at https://quoats.dill-picl.org/.
Open resource ↗pdf-page:1 lines:1-60The code for the web application discussed here is available at https://git.io/Jtv9J, and the application itself is available at https://quoats.dill-picl.org/.
Open resource ↗pdf-page:1 lines:1-60This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.