← Papers

Unverified paper record

Petal to the metal: The slow road to automating large-scale phenology labeling for herbarium specimens

bioRxiv · 22 Dec 2025 · 10.64898/2025.12.19.695552

Abstract

ABSTRACT Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found in-house expert-curated annotations were essential for producing reliable results. Expert validation found relatively strong accuracy for detecting present floral structures, but still had moderately high false negative rates. Applying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present. However, only 2.9 million of these contained complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts. Still, this dataset represents a large compilation of historical herbarium-derived phenology records available as a resource for the phenology community. We end by demonstrating how integrating these machine-labeled records into Phenobase, a publicly-available phenology database, expands taxonomic and temporal coverage for large-scale phenological analyses, and discuss remaining challenges and next steps.

Plant phenotyping relevance

植物標本画像から花の存在を自動検出し、精度検証と大規模な phenology データセット化を行う機械学習手法が研究の中心であるため、植物フェノタイピング手法として適格です。

abstractwe present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community.
abstractExpert validation found relatively strong accuracy for detecting present floral structures, but still had moderately high false negative rates.
abstractApplying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present.

Code and data availability

The paper's Data Availability Statement explicitly deposits the ensemble models, training/validation/test images, training data and final ensemble output on Zenodo, and the analysis code on GitHub; machine-labeled records are also served via the public Phenobase portal. All are paper-specific, public, and actionable.

Datasetpublic

ors contributed to drafts and gave final 454 approval for publication. 455 456 Data Availability Statement 457 The ensemble data models and a corresponding JSON file with model metadata data are 458 housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supportin

Open resource ↗Zenodo · 17675089 · pdf-raw-page:19 lines:1-55
Datasetpublic

tps://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supporting Information section at 466 the end of the article. 467 Appendix S1. List of difficult-to-annotate genera and families removed from training and 468 downstream data. 469 Appendix S2. Table S1. Validation results for held-o

Open resource ↗Zenodo · 10.5281/zenodo.17675089 · pdf-raw-page:19 lines:1-55
Codepublic

lity Statement 457 The ensemble data models and a corresponding JSON file with model metadata data are 458 housed on Zenodo (https://doi.org/10.5281/zenodo.17079402). Images used in training, 459 validation, and testing are located here: https://zenodo.org/records/17675089. Code used for 460 this project can be found on github (https://github.com/rafelafrance/phenobase/tree/v1.0.0).461 Training data and final ensemble output can be found on Zenodo 462 (https://doi.org/10.5281/zenodo.17675089).463 464 Supporting Information 465 Additional Supporting Information may be found online in the Supporting Information section at 466 the end of the article. 467 Appendix S1. List of difficult-to-annota

Open resource ↗GitHub · rafelafrance/phenobase · pdf-raw-page:19 lines:1-55

This is an automatically classified, unverified record. Curator approval is required before any resource enters the Catalog.