Unverified paper record
Genome-wide association mapping and predictive modeling of wet bean mass in a diverse cacao collection
bioRxiv · 26 Apr 2025 · 10.1101/2025.04.23.650203
Abstract
Improving cacao yield, a key objective in post-domestication crop improvement, remains a primary goal for breeders, but progress is often hindered by the confounding effects of population structure. To overcome this, we analyzed 346 diverse cacao accessions using an ML-based association mapping framework (with and without population structure adjustment) and a phenotype-only ML prediction of yield. By correcting for population structure, our Bootstrap Forest-based GWAS revealed association signals that showed consistent enrichment for ribosome and protein-synthesis functions, and a recurrent subset of SNPs with high importance appeared across multiple yield components, including pod index and seed number. In parallel, a Neural Network model was utilized to identify cotyledon mass and length as the most powerful predictors for total wet bean mass (R² = 0.715 by repeated five-fold cross-validation), suggesting a practical, low-cost screening proxy for breeding). Collectively, this study delivers a robust genetic framework and a novel predictive tool to accelerate the development of high-yielding cacao varieties through the early identification of elite clones.
Plant phenotyping relevance
カカオの形質(湿重量収量)を、測定可能な種子形質から予測するニューラルネットワークを開発・検証しており、低コストな表現型スクリーニング手法が研究の中心です。
abstracta phenotype-only ML prediction of yield
abstracta Neural Network model was utilized to identify cotyledon mass and length as the most powerful predictors for total wet bean mass (R² = 0.715 by repeated five-fold cross-validation)
abstractsuggesting a practical, low-cost screening proxy for breeding
Code and data availability
The supplied blocks describe GWAS and ML analyses of cacao yield, but the phenotype/genotype data are from a cited prior study (Bekele et al. 2022, ICGT database), not a paper-specific deposit by these authors. No author analysis code, scripts, trained models, or supplementary data with an explicit public availability/
No evidence-backed public reproduction asset is currently recorded.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.