This study introduces a two-step self-supervised machine learning framework for detecting biological assemblages in over 20 years of multifrequency water column sonar data archived by NOAA that have been converted into analysis-ready Zarr stores. In the first stage, only basic preprocessing is applied to ensure broad applicability across diverse sonar datasets. A shallow UNet then identifies clusters of interest in the unlabeled data, which are iteratively reviewed and annotated by domain experts in a continuous human-in-the-loop feedback process. In the second stage, these expert-validated pseudo labels are combined with existing prelabeled datasets to train a supervised fish classifier, overcoming the scarcity of labeled data while maintaining biological relevance. The novelty of this workflow lies in its scalability across large, heterogeneous archives with minimal preprocessing and in its seamless integration of unsupervised UNet clustering, expert validation, and downstream supervised fine tuning. By linking these components, the framework embodies self-supervised learning and delivers robust, accurate, and interpretable fish detection performance across decades of sonar surveys. A modular, open-source implementation using standard data formats ensures reproducibility, simplifies integration into existing pipelines, and invites community-driven extensions for diverse acoustic monitoring.
Hoelzemann et al. (Wed,) studied this question.