11 VCFs and metadata needed to reproduce Vaccinium sect. Cyanococcus and Ipomoea ser. 372 Batatas analyses with PPGTK will be made available via Dryad upon acceptance. PPGTK is 373 available on GitHub, and release v0.1.0-alpha was the version used for analyses in this 374 manuscript (https://github.com/tileylab/PPGTK/releases/tag/v0.1.0-alpha). PPGTK currently has 375 other functions for calculating population genetic summary statistics, but the classify-ploidy 376 function implements the machine-learning method described in the manuscript. 377 378 . CC-BY 4.0 International license is made available under a preprint (which was not certified by peer review) is the au
Open resource ↗tileylab/PPGTK · v0.1.0-alpha · pdf-raw-page:11 lines:1-24Unverified paper record
A simple and accurate method for inferring missing ploidy information from sequence data
10 Sept 2026 · 10.64898/2026.09.06.749761
Abstract
Polyploidy can be a critical factor for explaining plant trait variation, niche diversification, or speciation. However, inferring ploidy from silica-dried or historical samples using chromosome counts or flow cytometry is not possible, and scaling up ploidy estimation to population-level fresh contemporary samples can be challenging as well. Thus, we present a new method for estimating ploidy levels directly from sequencing data using machine learning; the Polyploid Population Genomics Tool Kit (PPGTK). The machine-learning approach is advantageous as it relaxes the assumptions of previous probabilistic methods and provides per-sample probabilities, allowing investigators to evaluate uncertainty in their system of interest.. We demonstrate performance and accuracy of the method on simulated and empirical data. Simulations showed above 99% accuracy, even for low coverage data, as long reads were mappable to the reference genome. For empirical analyses, we used target enrichment data from blueberry wild relatives (Vaccinium sect. Cyanococcus) and whole-genome data from sweetpotato wild relatives (Ipomoea ser. Batatas). Ploidy was recovered with 99% accuracy across 70 Vaccinium individuals and 97% across 82 Ipomoea individuals. Analysis of many individuals is fast and requires only a multisample VCF, which is presumably generated for the research anyway, and some samples of known ploidy for training the classifier. The approach implemented in PPGTK is promising for collections-based research as well, enabling ploidy classification of historical specimens based on present-day observations. The method is implemented in a new Python package as a single command that can run on a conventional laptop.
Plant phenotyping relevance
植物の倍数性という状態をシーケンスデータから推定する機械学習手法を開発し、シミュレーションおよび実データで精度検証している。Pythonパッケージとして実装され、手法自体が中心である。
abstractwe present a new method for estimating ploidy levels directly from sequencing data using machine learning; the Polyploid Population Genomics Tool Kit (PPGTK).
abstractWe demonstrate performance and accuracy of the method on simulated and empirical data.
abstractThe method is implemented in a new Python package as a single command that can run on a conventional laptop.
Code and data availability
The paper's ploidy-classification method is implemented in the authors' public Python package PPGTK, with a specific release (v0.1.0-alpha) used for the manuscript's analyses. The empirical VCF/metadata datasets are promised on Dryad only 'upon acceptance' and thus are not yet actionable.
This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.