Premise Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Methods Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found that in‐house expert‐curated annotations were essential for producing reliable results. Results Expert validation found that the modeling pipeline detected present floral structures with relatively strong accuracy but had moderately high false negative rates. Applying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present. However, only 2.9 million of these contained the complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts. Discussion The dataset presented here represents a large compilation of historical herbarium‐derived phenology records available as a resource for the phenology community. We end by demonstrating how integrating these machine‐labeled records into Phenobase, a publicly available phenology database, expands taxonomic and temporal coverage for large‐scale phenological analyses.
No takes yet. Share an insight, caveat, or question.
Grady et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: