Accurate registration of optical and synthetic aperture radar (SAR) images is a fundamental prerequisite for multi-source remote sensing data fusion and analysis. However, due to the substantial differences in imaging mechanisms, optical–SAR image pairs often exhibit significant radiometric discrepancies and spatially varying geometric inconsistencies, which severely limit the robustness of traditional feature or region-based registration methods in cross-modal scenarios. To address these challenges, this paper proposes an end-to-end Optical–SAR Registration Network (OSR-Net) based on multi-constraint joint optimization. The proposed framework explicitly decouples cross-modal feature alignment and geometric correction, enabling robust registration under large appearance variation. Specifically, a multi-modal feature extraction module constructs a shared high-level representation, while a multi-scale channel attention mechanism adaptively enhances cross-modal feature consistency. A multi-scale affine transformation prediction module provides a coarse-to-fine geometric initialization, which stabilizes parameter estimation under complex imaging conditions. Furthermore, an improved spatial transformer network is introduced to perform structure-preserving geometric refinement, mitigating spatial distortion induced by modality discrepancies. In addition, a multi-constraint loss formulation is designed to jointly enforce geometric accuracy, structural consistency, and physical plausibility. By employing a dynamic weighting strategy, the optimization process progressively shifts from global alignment to local structural refinement, effectively preventing degenerate solutions and improving robustness. Extensive experiments on public optical–SAR datasets demonstrate that the proposed method achieves accurate and stable registration across diverse scenes, providing a reliable geometric foundation for subsequent multi-source remote sensing data fusion.
Zhang et al. (Mon,) studied this question.