In photovoltaic remote sensing image segmentation tasks, fully supervised methods can achieve high accuracy. However, the high cost of pixel-level annotation significantly limits their scalability in large-scale scenarios. To overcome this annotation bottleneck, this paper proposes a zero-shot cross-modal segmentation framework based on the visual-language pre-trained foundation model (CLIP). This approach harnesses CLIP’s cross-modal knowledge transfer capabilities to achieve precise extraction of photovoltaic targets without requiring any downstream training. This paper first introduces the Layer-wise Augmented Residual Attention (LARA) mechanism to enhance fine-grained detail representation in the feature space. Subsequently, a Cross-modal Semantic Attribution Module (CMSA) is designed to generate precise activation maps by leveraging image-text alignment gradient information. Finally, the Confidence-Aware Refinement Strategy (CARS) replaces the conventional training-based denoising process, directly producing high-quality binary segmentation masks through adaptive thresholding. Comparative experiments were conducted to evaluate the proposed method against various baselines using several public datasets with varying resolutions in Jiangsu Province including Unmanned Aerial Vehicles imagery, Beijing-2, Gaofen-2, and a self-created Sentinel-2 imagery covering multiple countries. Notably, the proposed method achieved an IoU of 70. 3% on the Gaofen-2 PV03 dataset with a spatial resolution of approximately 0. 3 m and 50. 8% on the self-created Sentinel-2 PVSentinel-2 dataset with a spatial resolution of 10 m. Experimental results demonstrate that our proposed approach maintains excellent cross-domain generalisation capabilities while reducing annotation costs, thereby providing an efficient and viable technical pathway for the automated monitoring of large-scale photovoltaic facilities.
Li et al. (2026) studied this question.