• Repurposed the DCLDE 2024 dataset from localization to detection and classification of whale vocalizations. • Comparative evaluation of three deep learning pipelines: multiclass classification, detection-classification and binary relevance. • Proposed a prime distribution matrix to analyze inter-class behaviour in binary relevance predictions. • Conducted ablation studies on augmentation, uncertain samples and class imbalance. • Demonstrated monitoring applicability and cross-dataset generalization using DCLDE 2013 recordings. • Highlighted the impact of annotation quality and dataset design on bioacoustic monitoring performance. Tracking and monitoring whales is of high importance due to their many ecological roles. The vastness and inaccessibility of oceans necessitate automated solutions for long-term monitoring. In this work, the DCLDE 2024 workshop dataset, originally developed for localization tasks, is repurposed for detection and classification. Seven relevant annotated whale vocalizations and empty samples, are converted into images using three audio visualization techniques. We design and evaluate three deep learning pipelines for whale vocalization analysis: (1) full classification pipeline, (2) two-stage detection–classification pipeline and (3) binary relevance (one-vs-rest) pipeline. For the binary relevance approach, a novel prime distribution matrix is introduced to analyze inter-class behaviour and prediction patterns. Various ablation experiments are conducted to analyze the impact of data augmentation strategies, the inclusion of uncertain samples and class imbalance. Results show that non-augmented or lightly augmented datasets of confident samples outperform heavier augmentation, which can distort vocalization characteristics. The binary relevance pipeline achieves the strongest overall performance, reaching an F1-score of 62.46% on the full seven-class setup, while excluding two problematic classes improved the performance to 71.71%. Binary relevance also demonstrates strong applicability in continuous monitoring scenarios and generalization to new data sources. Using three classes from the DCLDE 2013 dataset, binary relevance achieves a macro-average F1-score of 67.99%. These results demonstrate the feasibility and flexibility of deep learning solutions for passive acoustic monitoring of whales. This study particularly highlights the importance of annotation quality and data variety, emphasizing the need for large and well-annotated datasets for future bioacoustic monitoring systems.
Blais et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: