Underwater salient instance segmentation (USIS) remains challenging because color distortion, low contrast, and scattering often weaken appearance cues, while reliable geometric measurements are usually unavailable. Existing methods mainly rely on red–green–blue (RGB) information, which can be insufficient in visually degraded underwater scenes. We propose the Depth-guided Cross-modal Fusion Network (DCFNet), a depth-guided fusion framework that leverages pseudo-depth estimated from RGB images as an auxiliary structural prior. DCFNet contains a dual-branch encoder, a cross-modal fusion branch, a refinement decoder, and an instance branch. In the fusion branch, the proposed Depth-Aware Modality Injection (DAMI) module selectively exchanges information between RGB and pseudo-depth features to reduce the influence of noisy depth estimates. The decoder further combines Inverted Residual Transformer (IRT) blocks and Bidirectional Attention Gate (BiAG) modules for contextual modeling and boundary refinement. Finally, the instance branch integrates positional cues to generate dynamic kernels for proposal-free mask prediction. Experiments on USIS10K and USIS16K show that DCFNet achieves competitive performance against several relevant baselines. Ablation studies further indicate that both the pseudo-depth prior and the proposed fusion architecture contribute to the final performance.
Zheng et al. (Thu,) studied this question.