The datasets and code supporting the conclusions of this article are available in the Zenodo repository, DOI: 10.5281/zenodo.15877862.
Open resource ↗Zenodo · 10.5281/zenodo.15877862 · lines:256-273Unverified paper record
When to cluster phenotypic data? A simulation-based framework to guide decisions in agrobiodiversity research
9 Jan 2026 · 10.21203/rs.3.rs-8551451/v1
Abstract
Abstract Phenotypic clustering is a cornerstone of population structure analysis in agrobiodiversity research, especially for neglected and underutilized species (NUS) where genomic data are scarce. However, there is currently no formal method to determine whether a given dataset contains sufficient biological signal to justify clustering, leading to potential overinterpretation of spurious patterns. To address this, we introduce a signal-first diagnostic framework. This framework mandates the assessment of phenotypic differentiation prior to any unsupervised classification, providing clear, data-driven thresholds to decide if clustering is statistically meaningful. We developed this framework through a large-scale, empirically-grounded simulation study. Using realistic trait architectures calibrated on fonio ( Digitaria exilis ), we evaluated 11 clustering algorithms across a continuous gradient of phenotypic differentiation (Pst = 0.05–0.85). Our results establish quantitative detectability thresholds: under the calibrated trait architecture, clustering fails to recover meaningful structure below Pst ≈ 0.30, a range typical for many NUS. Even the best-performing algorithm required Pst > 0.47 for moderate accuracy. We further demonstrate that internal validation metrics (e.g., Silhouette score) are unreliable under weak differentiation, often misleadingly suggesting robust clusters. The proposed framework shifts the analytical paradigm from algorithm selection to signal assessment. We provide practical guidelines and an openly available simulation template to help researchers implement this workflow, thereby supporting more reliable diversity assessments, core collection design, and germplasm management decisions in data-scarce systems.
Plant phenotyping relevance
植物の表現型データを対象に、クラスタリングの妥当性を事前評価する統計的診断フレームワークをシミュレーションで開発しており、再利用可能な表現型解析手法が中心である。
abstractTo address this, we introduce a signal-first diagnostic framework.
abstractWe developed this framework through a large-scale, empirically-grounded simulation study.
abstractWe provide practical guidelines and an openly available simulation template to help researchers implement this workflow
Code and data availability
The preprint states that simulation scripts, clustering implementations, parameter sets, and representative synthetic datasets are publicly available on Zenodo, with a specific DOI (10.5281/zenodo.15877862) given in the data availability statement. This is a paper-specific, publicly actionable asset covering the studyâ
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.