Key points are not available for this paper at this time.
Fault detection in power transmission line inspection plays a critical role in ensuring the stable operation of power grids and maintaining power supply quality. However, the visual characteristics of faults in inspection images, such as small sizes, partial occlusion and complex backgrounds, prevent existing deep-learning-based detectors from learning sufficient knowledge for accurate detection. To address this issue, we propose a novel fault detection method that utilizes additional multimodal knowledge as supplementation. A two-step training strategy is first designed to preliminarily introduce multimodal knowledge from transmission line image–text pairs into a Deformable DETR detector via a CLIP vision-language model, thereby overcoming the above challenge. Then, two customized improvements are proposed to further improve detection accuracy by fully leveraging the multimodal knowledge. To alleviate the degradation of multimodal knowledge during passing through Deformable DETR’s decoder, a knowledge-ensembled decoder block structure is presented to supplement multimodal knowledge through linear interpolation. Meanwhile, to ensure that the detector is optimized under the guidance of the multimodal knowledge, a pseudo-label distillation learning objective is introduced to provide an auxiliary supervision signal and achieve a better training effect. Experimental results show that our method achieves 4.4% and 5.6% enhancement in mAP@0.5 and mAP@0.5:0.95 compared to the baseline, while significantly outperforming existing transmission line fault detection methods. This proves the feasibility of improving fault detection performance by utilizing texts paired with images, which can be extended to numerous power system inspection scenarios.
Zhang et al. (Sat,) studied this question.