Key points are not available for this paper at this time.
Accurate meteorological visibility estimation is vital for transportation safety and environmental monitoring. However, modeling the inherent nonlinear spatial and spectral degradations in hazy environments remains challenging. While recent Large Vision-Language Models (LVLMs) offer strong scene understanding, they lack the regression precision required for visibility estimation. In this paper, we propose the Visibility-Aware Refined CNN (VISR-CNN), a dual-stream architecture that synthesizes local spatial cues with global frequency-domain signatures. The model integrates a Multi-Scale Transmission Attention (MSTA) module, which uses parallel dilated convolutions to estimate atmospheric transmission, and a Global Frequency Branch that utilizes 2D Real Fast Fourier Transforms (RFFT) with Spectral Gating to quantify visibility-dependent blurring. A progressive training strategy is introduced to decouple spectral and spatial optimization, and a physics-informed loss function is designed to supervise numerical regression while enforcing a monotonic ranking constraint consistent with physical light-attenuation laws. Results on the HKCHC-VD dataset show that VISR-CNN achieves state-of-the-art performance (MAE: 1.54 km; RMSE: 2.31 km), representing a 13.0% improvement over VisNet. Further evaluations on the CP1 and SWH datasets confirm robust generalization, reducing overall MAE by 21% and 20%, respectively, compared with the hybrid ResNeXt-50 + ViT model. Notably, in safety-critical range (0–10 km), VISR-CNN reduces RMSE for the HKCHC-VD, CP1, and SWH datasets by approximately 55%, 64%, and 71%, respectively, when compared with VisNet. These findings demonstrate the superiority of specialized, physics-grounded architectures over general-purpose LVLMs for high-precision meteorological regression.
Lo et al. (Thu,) studied this question.