← Papers

Unverified paper record

Empirically calibrated simulations reveal the limits of phenotypic clustering algorithms for biodiversity assessment in data-scarce crops.

PloS one · 17 Dec 2025 · 10.1371/journal.pone.0329254

Abstract

Clustering algorithms are widely used for phenotypic characterization and germplasm management, particularly in data-scarce crops such as neglected and underutilized species (NUS) that lack genomic resources. However, their performance under biologically realistic conditions remains poorly understood. Standard clustering methods commonly applied in crop research often assume distinct, isotropic, and homogeneous clusters, assumptions rarely satisfied in real-world phenotypic datasets. We developed a flexible and empirically calibrated simulation framework, using phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios. Our simulations integrated heterogeneous trait distributions (normal, gamma), strong inter-trait correlations (up to r = -0.84), heteroscedasticity, and moderate population structure (mean Pst = 0.16 ± 0.001, achieved through iterative calibration). Each scenario was replicated 100 times, with clustering accuracy evaluated using external (ARI, NMI) and internal (Silhouette, Davies-Bouldin) validation metrics under standardized conditions. The results revealed consistently poor algorithm performance under realistic conditions (e.g., ARI < 0.07), including for widely used methods in Neglected and Underutilized Species (NUS) research such as K-means, GMM, and PAM. Notably, conventional validation metrics failed to detect biologically meaningful structure revealed by geometric diagnostics, highlighting a critical methodological limitation. Performance markedly improved under idealized conditions, validating our simulation framework. These findings highlight the risk of overinterpreting clustering outputs from weakly structured phenotypic datasets and expose key limitations in current biodiversity analysis practices, particularly those guiding plant genetic resource conservation programs. We provide an open-source R-based diagnostic tool, with parameter specifications to assist practitioners in selecting reproducible and interpretable clustering approaches for germplasm management and biodiversity assessment in data-scarce crops.

Plant phenotyping relevance

植物の表現型データを対象に、クラスタリング手法を現実的な条件でベンチマークするシミュレーション枠組みとR診断ツールを開発しており、表現型解析手法が研究の中心である。

abstractWe developed a flexible and empirically calibrated simulation framework, using phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios.
abstractWe provide an open-source R-based diagnostic tool, with parameter specifications to assist practitioners in selecting reproducible and interpretable clustering approaches for germplasm management and biodiversity assessment in data-scarce crops.

Code and data availability

The paper's Data Availability statement explicitly deposits the complete R simulation/clustering/evaluation script on Zenodo (DOI 10.5281/zenodo.15877863), a paper-specific, publicly actionable code asset. The empirical fonio trait data belong to a prior cited study (Bio et al.), not this paper, and supporting files/DO

Codepublic

the complete R script used to simulate phenotypic datasets, apply clustering algorithms, and compute evaluation metrics is publicly available on Zenodo: https://doi.org/10.5281/zenodo.15877863

Open resource ↗Zenodo · 10.5281/zenodo.15877863 · lines:107-122

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.