• Markov process design adapted for dynamic satellite mission environments • Adaptive reward mechanism for dynamic mission scheduling • Meta-RL for hybrid mission generalization Earth observation satellites (EOSs) play crucial roles in disaster monitoring, resource management, military reconnaissance, and environmental protection. However, the increasing complexity and dynamic nature of EOS missions pose significant challenges for conventional mission scheduling methods, which frequently struggle with high computational overhead, poor adaptability to real-time changes, and limited generalizability across mission scenarios. Thus, this paper proposed a meta reinforcement learning (MRL) framework for hybrid dynamic mission scheduling (DMS) for EOSs. The proposed MRL-DMS method integrates a mission-adaptive reward mechanism and a dynamic action space within a Markov decision process formulation to enable real-time responsiveness to stochastic mission arrivals. A metalearning layer based on long short-term memory networks is implemented to enhance policy generalization across diverse mission distributions. Leveraging proximal policy optimization (PPO) as the inner loop reinforcement learning algorithm, the MRL-DMS method adapts to new mission environments with minimal retraining. The proposed method was evaluated on realistic satellite orbit data and largescale hybrid mission datasets derived from global conflict scenarios. The experimental results demonstrate that the MRL-DMS method outperforms the state-of-the-art PPO, A2C, and DQN algorithms. The MRL-DMS method achieves significant improvements in the dynamic mission completion rates, cumulative reward acquisition, and computational efficiency. In addition, MRL-DMS exhibits robust adaptability across varying mission densities, scales, and spatial distributions, effectively prioritizing high-urgency missions while maintaining stable performance under scheduling pressure. The MRL-DMS method provides a scalable, intelligent solution for autonomous satellite mission planning in dynamic and uncertain operational environments. The findings of this study indicate that the MRL-DMS method provides valuable insights into real-time remote sensing strategies in response to global emergencies and rapidly evolving geopolitical events.
Yao et al. (2026) studied this question.