As the volume of active water column sonar data expands, automated, machine learning (ML)-based echogram analysis methods are increasingly used to detect biological scatterers at scale. Semantic segmentation, which labels each pixel in an image, is a class of methods that is increasingly adopted to detect and classify acoustic targets. However, survey echograms are not standard images: they have spatiotemporal associations and originate from data collection strategies targeting potentially patchy biological aggregations. In the context of transect surveys, we consider four key requirements to create echogram datasets for ML applications: (1) how to partition data based on survey information, (2) how to subsample spatially via transect-based groupings to minimize model overfit, (3) how to create region masks of biological scatterers to exclude noise and other contaminating sources (e.g., seafloor), and (4) how to reconcile different spatiotemporal resolutions of echo data and human annotations. We present a generalizable workflow for constructing echogram datasets that addresses these requirements, and discuss our implementation using two open-source tools, Echopype and Echoregions. We highlight how the workflow enables a flow of echo data with ancillary spatiotemporal information propagating downstream, and demonstrate its scalability in organizing a multi-year dataset from a fisheries acoustic-trawl survey.
Tuguinay et al. (Tue,) studied this question.