In recent years, Deep Reinforcement Learning (DRL) models have demonstrated potential in effectively addressing the Agile Earth Observation Satellite Scheduling Problem (AEOSSP). However, these policy models prioritize optimizing overall policy expectations over individual decision accuracy, resulting in decision errors in certain scenes. To mitigate this issue, we propose a DRL-based Self-Repair Construction Method (SRCM), which is a two-stage method that includes an Improved Construction Model (ICM) and a Self-Repair Process (SRP). The ICM, an encoder-decoder based neural policy model, is specifically designed to construct initial solutions for the AEOSSP. The SRP incorporates two mechanisms relaxation-insertion and relaxation-masking to investigate and repair decision-making errors in DRL model solutions. Comparative experiments demonstrate that the proposed SRCM surpasses the state-of-the-art problem-specific meta-heuristics in both solution quality and timeliness. The results from our model study indicate that the ICM within SRCM outperforms other neural policy models in extensive validations. Moreover, the mechanism study of the SRP shows that decision-making errors are prevalent across all test neural policy models and confirms the effectiveness of the SRP in rectifications.
Publications
- Article type
- Year
- Co-author
Year
Open Access
Issue
Tsinghua Science and Technology 2026, 31(1): 180-198
Published: 25 August 2025
Downloads:270
Total 1
京公网安备11010802044758号