Accurate crack segmentation is critical for infrastructure monitoring but remains challenging due to diverse crack morphologies, complex backgrounds, and domain shifts. This paper proposes DMSRCrack, a Dual-encoder Multi-Scale Refinement network for robust crack segmentation across diverse domains. The proposed architecture integrates a hybrid CNN–ViT encoder for local and global feature extraction, a Crack Detail Enhancement Module (CDEM) for preserving thin crack structures, a Boundary Refinement Head (BRH) for contour sharpening, and a Multi-Scale Fusion (MSF) module for scale-consistent representation. Across eight benchmark datasets, DMSRCrack achieves an average Dice of 0.7949 ± 0.0655 and IoU of 0.6639 ± 0.0917 in dataset-wise training. Under leave-one-dataset-out evaluation, it attains the highest IoU on DeepCrack (0.4056), Rissbilder (0.3540), and Crack500 (0.3008). Ablation and computational analyses further confirm the effectiveness and practical efficiency of the proposed contributions. • Dual-encoder network captures local texture and global crack structure. • CDEM, BRH, and MSF modules preserve thin cracks and sharpen boundaries. • Composite loss improves pixel accuracy and topological consistency. • Strong cross-dataset performance under leave-one-dataset-out testing. • Favorable accuracy-efficiency balance with 11.66M parameters.
Al-Sameai et al. (2026) studied this question.