To achieve pixel‐level semantic segmentation of surface cracks in cultural relic buildings, depth‐separable convolution (DSC), neighboring information fusion (NIF) module, dual‐domain coordination module (DCM), and feature refinement module (FRM) are introduced into the U‐Net model. Construct an improved U‐Net model (MAU‐Net) with multiscale perception optimization suitable for crack identification in cultural relic buildings. DSC reconstructs the decoder to reduce the computational load of the model. NIF effectively fuses the feature information of adjacent layers. DCM integrates the characteristics of the Convolutional Block Attention Module (CBAM) and the Channel Squeeze and Excitation (CSE). FRM further optimizes the extracted features to improve the accuracy and robustness of segmentation. Experiments were conducted on the self‐made dataset of gaps on the Ming and Qing Dynasty city walls in Kaifeng. MAU‐Net was compared against several state‐of‐the‐art models, including U‐Net, DcsNet, SegFormer, CrackformerII, DECS, DTrc, TransMUNet, and DeepLabv3+. The results show that the Pr, Re, F1, and MIoU of the MAU‐Net model are 87.36%, 77.98%, 82.41%, and 72.53%, respectively, which are higher than those of other models. The comparison of the performance of the detection tasks, the classification confusion matrix, and the heat map generated by the Grad‐CAM method shows that the MAU‐Net model has the best detection effect. The research results provide a high‐precision and lightweight automated method for the health monitoring of brick cultural relic buildings.
Zhang et al. (Thu,) studied this question.