← Papers

Unverified paper record

A gene-phenotype relationship extraction pipeline from the biomedical literature using a representation learning approach.

Bioinformatics (Oxford, England) · 1 Jul 2018 · 10.1093/bioinformatics/bty263

Abstract

Motivation The fundamental challenge of modern genetic analysis is to establish gene-phenotype correlations that are often found in the large-scale publications. Because lexical features of gene are relatively regular in text, the main challenge of these relation extraction is phenotype recognition. Due to phenotypic descriptions are often study- or author-specific, few lexicon can be used to effectively identify the entire phenotypic expressions in text, especially for plants. Results We have proposed a pipeline for extracting phenotype, gene and their relations from biomedical literature. Combined with abbreviation revision and sentence template extraction, we improved the unsupervised word-embedding-to-sentence-embedding cascaded approach as representation learning to recognize the various broad phenotypic information in literature. In addition, the dictionary- and rule-based method was applied for gene recognition. Finally, we integrated one of famous information extraction system OLLIE to identify gene-phenotype relations. To demonstrate the applicability of the pipeline, we established two types of comparison experiment using model organism Arabidopsis thaliana. In the comparison of state-of-the-art baselines, our approach obtained the best performance (F1-Measure of 66.83%). We also applied the pipeline to 481 full-articles from TAIR gene-phenotype manual relationship dataset to prove the validity. The results showed that our proposed pipeline can cover 70.94% of the original dataset and add 373 new relations to expand it. Availability and implementation The source code is available at http://www.wutbiolab.cn: 82/Gene-Phenotype-Relation-Extraction-Pipeline.zip. Supplementary information Supplementary data are available at Bioinformatics online.

Plant phenotyping relevance

植物の表現型情報を文献から認識・抽出し、遺伝子との関係を構築する計算パイプライン自体が研究の中心であり、Arabidopsisで性能評価も行っている。

abstractWe have proposed a pipeline for extracting phenotype, gene and their relations from biomedical literature.
abstractour approach obtained the best performance (F1-Measure of 66.83%).
abstractwe established two types of comparison experiment using model organism Arabidopsis thaliana.

Code and data availability

The paper's authors explicitly state that the source code for their gene–phenotype relationship extraction pipeline (the computational analysis of this paper) is publicly available for download at their lab's URL (wutbiolab.cn). No paper-specific phenotype datasets, images, or trained model checkpoints are described as

Codepublic

The source code is available at http://www.wutbiolab.cn: 82/Gene-Phenotype-Relation-Extraction-Pipeline.zip .

Open resource ↗Gene-Phenotype-Relation-Extraction-Pipeline · lines:1-33

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.