The scheduling of Earth observation satellites presents a formidable multi-objective optimization challenge, characterized by inherent trade-offs among task completion rate, execution timeliness, and the temporal uniformity of revisits. To address this, we introduce the Multi-Satellite Observation Task Scheduling (MSOTS) framework, a novel end-to-end approach based on Multi-Agent Reinforcement Learning (MARL). This framework formulates the scheduling process as a Markov game, employing the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm within a Centralized Training, Decentralized Execution (CTDE) paradigm to effectively navigate these competing objectives. Furthermore, to ensure a balanced evaluation, we propose a Composite Multi-Objective Performance Score grounded in a weighted harmonic mean. Comprehensive empirical evaluations conducted on large-scale, simulated orbital scenarios demonstrate that MSOTS significantly outperforms both traditional heuristics and existing deep reinforcement learning methods in comprehensive performance and robust efficiency. This research provides a highly effective and intelligent approach to modern satellite task scheduling.
Zhang et al. (2026) studied this question.