← Papers

Unverified paper record

DIVIS: a semantic DIstance to improve the VISualisation of heterogeneous phenotypic datasets.

BioData mining · 4 Apr 2022 · 10.1186/s13040-022-00293-y

Abstract

Background Thanks to the wider spread of high-throughput experimental techniques, biologists are accumulating large amounts of datasets which often mix quantitative and qualitative variables and are not always complete, in particular when they regard phenotypic traits. In order to get a first insight into these datasets and reduce the data matrices size scientists often rely on multivariate analysis techniques. However such approaches are not always easily practicable in particular when faced with mixed datasets. Moreover displaying large numbers of individuals leads to cluttered visualisations which are difficult to interpret. Results We introduced a new methodology to overcome these limits. Its main feature is a new semantic distance tailored for both quantitative and qualitative variables which allows for a realistic representation of the relationships between individuals (phenotypic descriptions in our case). This semantic distance is based on ontologies which are engineered to represent real-life knowledge regarding the underlying variables. For easier handling by biologists, we incorporated its use into a complete tool, from raw data file to visualisation. Following the distance calculation, the next steps performed by the tool consist in (i) grouping similar individuals, (ii) representing each group by emblematic individuals we call archetypes and (iii) building sparse visualisations based on these archetypes. Our approach was implemented as a Python pipeline and applied to a rosebush dataset including passport and phenotypic data. Conclusions The introduction of our new semantic distance and of the archetype concept allowed us to build a comprehensive representation of an incomplete dataset characterised by a large proportion of qualitative data. The methodology described here could have wider use beyond information characterizing organisms or species and beyond plant science. Indeed we could apply the same approach to any mixed dataset.

Plant phenotyping relevance

植物の表現型データを対象に、混合型・不完全データを可視化する新しい意味距離とPythonパイプラインを開発しており、表現型解析手法が研究の中心である。

abstractWe introduced a new methodology to overcome these limits.
abstractThis semantic distance is based on ontologies which are engineered to represent real-life knowledge regarding the underlying variables.
abstractOur approach was implemented as a Python pipeline and applied to a rosebush dataset including passport and phenotypic data.

Code and data availability

The authors' DIVIS Python pipeline (semantic distance, clustering, archetype visualisation) is publicly available on Forgemia with explicit availability language; the rosebush phenotype dataset itself is only available on request, so it is not a qualifying public asset.

Codepublic

Loire”, supported by the French Region Pays de la Loire, Angers Loire Métropole and the European Regional Development Fund, as part of the DIVIS project. Availability of data and materials The software developed to implement the pipeline presented in this paper is available as follows: • Project name: DIVIS • Project home page: https://forgemia.inra.fr/irhs-bioinfo/Divis • Archived version: v1.2 • Operating system(s): Platform independent • Programming language: Python 3.7 • Other requirements: Described as requirement.txt file for pip in the code repository • License: CeCILL. See LICENCE file in the code repository • Any restrictions to use by non-academics: None The OWL ontology (in French

Open resource ↗irhs-bioinfo/Divis · v1.2 · lines:431-490

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.