• Fuses vibration imaging and deep learning with D-S theory for gear pitting diagnosis. • Heterogeneous fusion of vibration and image data overcomes single-source limitations. • Achieves 97.26% accuracy on custom test bench, outperforming single-modality methods. As critical transmission components, gears are highly susceptible to pitting damage, the accurate identification of which is essential for operational safety and maintenance efficiency. Conventional vibration analysis often yields ambiguous fault signatures, while visual inspection is prone to interference under varying operating conditions. To overcome the limitations of single-modal approaches, a dedicated test bench was constructed to synchronously acquire vibration and image data. Based on this dual-modality setup, we propose a deep learning framework for effectively fusing heterogeneous data to achieve precise fault diagnosis. Multiple vibration visualization techniques, including Symmetrical Dot Pattern (SDP), Recurrence Plot (RP), and time-frequency representations, were systematically evaluated using networks such as VGG, AlexNet, and Swin-Transformer (S-T) to determine the optimal combination. Comparative analysis identified SDP processed by VGG16 and S-T as the most effective approach, which was subsequently adopted to generate vibration evidence. For image data, augmentation was performed using DCGAN, followed by gear tooth extraction, tilt and distortion correction, and precise segmentation with an improved U²-Net model to extract visual evidence. A state recognition network was then established to integrate these heterogeneous sources, with decision-level fusion implemented via D-S evidence theory. Experimental results demonstrate that the proposed fusion method effectively combines multi-modal information, achieving a pitting identification accuracy of 97.26%, outperforming vibration-only and image-only methods by 5.9% and 3.7%, respectively. The framework thus offers a more reliable and comprehensive solution for gear pitting detection.
Gu et al. (2026) studied this question.