Background Passive acoustic monitoring (PAM) combined with deep learning is increasingly used to detect cetacean vocalizations. The workflow from raw audio to a trained detector (annotation, segmentation, spectrogram parameterization, image processing, and model evaluation) is rarely documented in sufficient detail to allow other groups to reproduce or compare results. This lack of standardization limits methodological comparability across studies and hinders systematic investigation of how upstream preprocessing choices affect detection performance. Methods This article presents ai-pam-pipeline, an open-source six-stage framework that formalizes the complete CNN-based detection workflow, with every parameter defined in a single YAML configuration file ensuring exact experimental reproducibility. The framework separates the methodological structure from its computational implementation, allowing adaptation to different species, vocalization types, or acoustic environments, through configuration alone, without code modification. Validation was conducted on recordings of bottlenose dolphin (Tursiops truncatus) whistles through two complementary experiments: (A) a binary whistle detector trained with controlled environment recordings, and evaluated on an independent open-ocean dataset, and (B) a five-class vocalization classifier demonstrating framework scalability. Results The binary detector achieved macro F1 = 0.966 under ten-folds cross-validation on in-domain data. Cross-domain evaluation produced zero false positives across all ten folds (false discovery rate = 0.000) with macro F1 = 0.759, confirming generalization to unseen acoustic conditions without retraining. The multiclass extension achieved macro F1 = 0.843 across five vocalization categories; inter-class confusion between echolocation click trains and burst-pulse sounds reflects biological signal overlap rather than classifier failure. Conclusions The framework provides a reproducible and documented baseline for CNN-based cetacean detection in PAM. The complete toolkit is publicly released under the Apache 2.0 license, while the datasets used in this study are made openly available through separate public repositories. Together, these resources enable exact replication of the analyses and support adaptation to new monitoring contexts.
Rocco De Marco (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: