Remote sensing image change captioning (RSICC) aims to describe changes between bi-temporal remote sensing (RS) images in natural language. The existing methods, typically based on traditional encoder–decoder architectures or multi-task collaboration, face limitations in terms of either their description accuracy or computational efficiency. To address these challenges, we propose the MTH-Net, a Mamba–Transformer Hybrid Network with joint spatiotemporal awareness. The model is built upon a symmetric Siamese network to extract comparable features from the input image pair. The MTH-Net introduces a multi-class feature generation (MFG) module to produce diverse features tailored for spatiotemporal modeling. Its core Spatiotemporal Difference Perception Network (SDPN) effectively integrates Mamba for efficient long-sequence temporal modeling and Transformer for fine-grained spatial dependency capture, leveraging a broadcasting mechanism for complementary fusion. A feature-sharing strategy is employed to reduce the computational overhead in multi-task learning. Extensive experiments on the LEVIR-CDC, WHU-CDC, and LEVIR-MCI datasets demonstrate that the MTH-Net achieves a state-of-the-art performance in change captioning, validating the effectiveness of our hybrid design and feature-sharing mechanism.
Ma et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: