Underwater ship recognition is fundamentally challenged by acoustic complexity and limited labeled data. This paper proposes a physics-driven multidimensional feature fusion (MDFF) framework via auditory-inspired cross-attention learning. Rather than relying on isolated acoustic features or simple shallow concatenation, the method integrates bio-inspired Gammatone filtering with both LOFAR (structural line-spectra) and DEMON (propeller modulation) representations. A deep cross-attention architecture (MDFF-SL) is designed to effectively decouple machinery and propulsion signatures, mapping these physically heterogeneous components into a compact, highly discriminative representation space while suppressing overfitting under small-sample conditions. Extensive experiments on real-world sea-trial data demonstrate that MDFF-SL achieves 92.96% recognition accuracy with exceptionally low standard deviation, substantially outperforming conventional single-domain and shallow-fusion baselines. These results establish a high-precision, computationally efficient paradigm for reliable maritime surveillance in non-stationary underwater environments.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: