← Papers

Unverified paper record

When to cluster phenotypic data? A simulation-based framework to guide decisions in agrobiodiversity research

9 Jan 2026 · 10.21203/rs.3.rs-8551451/v1

Abstract

Abstract Phenotypic clustering is a cornerstone of population structure analysis in agrobiodiversity research, especially for neglected and underutilized species (NUS) where genomic data are scarce. However, there is currently no formal method to determine whether a given dataset contains sufficient biological signal to justify clustering, leading to potential overinterpretation of spurious patterns. To address this, we introduce a signal-first diagnostic framework. This framework mandates the assessment of phenotypic differentiation prior to any unsupervised classification, providing clear, data-driven thresholds to decide if clustering is statistically meaningful. We developed this framework through a large-scale, empirically-grounded simulation study. Using realistic trait architectures calibrated on fonio ( Digitaria exilis ), we evaluated 11 clustering algorithms across a continuous gradient of phenotypic differentiation (Pst = 0.05–0.85). Our results establish quantitative detectability thresholds: under the calibrated trait architecture, clustering fails to recover meaningful structure below Pst ≈ 0.30, a range typical for many NUS. Even the best-performing algorithm required Pst > 0.47 for moderate accuracy. We further demonstrate that internal validation metrics (e.g., Silhouette score) are unreliable under weak differentiation, often misleadingly suggesting robust clusters. The proposed framework shifts the analytical paradigm from algorithm selection to signal assessment. We provide practical guidelines and an openly available simulation template to help researchers implement this workflow, thereby supporting more reliable diversity assessments, core collection design, and germplasm management decisions in data-scarce systems.

Plant phenotyping relevance

植物の表現型データを対象に、クラスタリングの妥当性を事前評価する統計的診断フレームワークをシミュレーションで開発しており、再利用可能な表現型解析手法が中心である。

abstractTo address this, we introduce a signal-first diagnostic framework.
abstractWe developed this framework through a large-scale, empirically-grounded simulation study.
abstractWe provide practical guidelines and an openly available simulation template to help researchers implement this workflow

Code and data availability

The preprint states that simulation scripts, clustering implementations, parameter sets, and representative synthetic datasets are publicly available on Zenodo, with a specific DOI (10.5281/zenodo.15877862) given in the data availability statement. This is a paper-specific, publicly actionable asset covering the studyâ

Codepublic

The datasets and code supporting the conclusions of this article are available in the Zenodo repository, DOI: 10.5281/zenodo.15877862.

Open resource ↗Zenodo · 10.5281/zenodo.15877862 · lines:256-273

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.