← Papers

Unverified paper record

DIVIS: A Semantic Distance to Improve the Visualization of Incomplete Heterogeneous Phenotypic Datasets

2 Aug 2021 · 10.21203/rs.3.rs-742853/v1

Abstract

Abstract BackgroundThanks to the wider spread of high-throughput experimental techniques, biologists are accumulating large amounts of datasets which often mix quantitative and qualitative variables and are not always complete, in particular when they regard phenotypic traits. In order to get a first insight into these datasets and reduce the data matrices size scientists often rely on multivariate analyses. However such approaches are not always easily practicable in particular when faced with mixed datasets with missing values. Moreover displaying large numbers of individuals leads to cluttered visualizations which are difficult to interpret. ResultsWe introduce a new methodology to overcome these limits. The underlying principle consists in (i) grouping similar individuals, (ii) representing each group by emblematic individuals we call archetypes and (iii) build sparse visualizations based on these archetypes. As a preliminary step to the clustering we design a new semantic distance tailored for both quantitative and qualitative variables which allows a realistic representation of the relationships between individuals. This semantic distance is based on ontologies which are engineered to represent real life knowledge regarding the underlying variables. Our approach is implemented as a Python pipeline and illustrated by a rosebush dataset including passport and phenotypic data. ConclusionsThe introduction of our new semantic distance and of the archetype concept allows us to build a comprehensive representation of an incomplete dataset characterized by large proportion of qualitative data. The methodology described here could have wider use beyond information characterizing organisms or species and beyond plant science. Indeed we could apply the same approach to any incomplete mixed dataset.

Plant phenotyping relevance

不完全で異種の植物表現型データを可視化するための距離尺度・クラスタリング・アーキタイプ表現・Pythonパイプラインを中心に開発しており、表現型解析手法が研究の中核である。

abstractWe introduce a new methodology to overcome these limits.
abstractAs a preliminary step to the clustering we design a new semantic distance tailored for both quantitative and qualitative variables
abstractOur approach is implemented as a Python pipeline and illustrated by a rosebush dataset including passport and phenotypic data.

Code and data availability

The authors' DIVIS Python pipeline (semantic distance, clustering, archetype visualization) is publicly available on Forgemia with an archived v1.0 release and bundled OWL ontology. The rosebush phenotypic dataset itself is only available on request, so it is not a public asset.

Codepublic

is • MCA: Multiple Correspondance Analysis • MDS: Multi-Dimensional Scaling • OWL: Web Ontology Language • SPARQL: SPARQL Protocol and RDF Query Language Availability of data and materials The software developed to implement the pipeline presented in this paper is available as follows: • Project name: DIVIS • Project home page: https://forgemia.inra.fr/irhs-bioinfo/Divis • Archived version: v1.0 • Operating system(s): Platform independent • Programming language: Python 3.7 • Other requirements: Described as requirement.txt file for pip in the code repository • License: CeCILL. See LICENCE file in the code repository • Any restrictions to use by non-academics: None The OWL ontology (in French

Open resource ↗irhs-bioinfo/Divis · v1.0 · pdf-raw-page:17 lines:1-69

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.