PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026The Journal of the Acoustical Society of America0 citations

Comparing deep learning architectures for call type identification in baleen whale repertoires with few categories

View Full Paper
GAGabriela C. AlongiNational Marine Mammal FoundationEFElizabeth FergusonUnited States Army Corps of EngineersPSPeter SugarmanBellevue College

Key Points

  • This study aims to evaluate different deep learning architectures for identifying baleen whale vocalizations.
  • Utilized the DeepAcoustics tool to analyze two CNN architectures: TinyYOLO (lightweight) and DarkNet (heavyweight).
  • Both networks were pre-trained on the COCO dataset and adapted for spectrogram-based classification.
  • Employed long-term acoustic recordings from Antarctica and the North Pacific to assess call identification between species.
  • TinyYOLO and DarkNet demonstrated varying levels of effectiveness in identifying specific whale call types.
  • Challenges in accurate categorization highlighted, particularly related to network architecture selection and customization.
  • The use of bounding-box detections showed potential for broader acoustic category grouping.

Abstract

Deep learning methods are increasingly applied to categorize animal vocalizations and have the potential to contribute to the complex task of determining structured call repertoires. In this study, we used the DeepAcoustics tool to evaluate two convolutional neural network architectures—TinyYOLO (lightweight) and DarkNet (heavyweight)—for multiclass detection of predefined baleen whale call types in long-term acoustic recordings from Antarctica and the North Pacific. Both networks were pre-trained on the COCO imagery dataset and adapted for spectrogram-based classification. We assessed how effectively each network identifies specific call types within a constrained repertoire and between species, including blue whale A, B, D, and Z calls, as well as 20- and 40-Hz fin whale calls. We also examined the potential of using bounding-box detections to group calls into broader acoustic categories—such as downsweeps or tonal units—as a means of coarse repertoire assessment. Our findings highlight strengths and limitations in applying object detection to marine bioacoustics and emphasize the importance of network architecture selection and customization in supporting accurate call-type categorization and comparative acoustic analysis across and within species.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Alongi et al. (2025) studied this question.

synapsesocial.com/papers/6a05684ea550a87e60a20c87https://doi.org/10.1121/10.0040871
Ask AI
Helpful
Bookmark
Share
View Full Paper