AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (4.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Open Access

A Deep Reinforcement Learning-Based Self-Repair Method for Solving the Agile Satellite Scheduling Problem

College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
Communications Technology Institute, Cairo 11765, Egypt
Shandong Provincial Institute of Land Surveying and Mapping, Jinan 250013, China

Yahui Zuo and Ming Chen contribute equally to this paper.

Show Author Information

Abstract

In recent years, Deep Reinforcement Learning (DRL) models have demonstrated potential in effectively addressing the Agile Earth Observation Satellite Scheduling Problem (AEOSSP). However, these policy models prioritize optimizing overall policy expectations over individual decision accuracy, resulting in decision errors in certain scenes. To mitigate this issue, we propose a DRL-based Self-Repair Construction Method (SRCM), which is a two-stage method that includes an Improved Construction Model (ICM) and a Self-Repair Process (SRP). The ICM, an encoder-decoder based neural policy model, is specifically designed to construct initial solutions for the AEOSSP. The SRP incorporates two mechanisms relaxation-insertion and relaxation-masking to investigate and repair decision-making errors in DRL model solutions. Comparative experiments demonstrate that the proposed SRCM surpasses the state-of-the-art problem-specific meta-heuristics in both solution quality and timeliness. The results from our model study indicate that the ICM within SRCM outperforms other neural policy models in extensive validations. Moreover, the mechanism study of the SRP shows that decision-making errors are prevalent across all test neural policy models and confirms the effectiveness of the SRP in rectifications.

References

【1】
【1】
 
 
Tsinghua Science and Technology
Pages 180-198

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zuo Y, Chen M, Liu X, et al. A Deep Reinforcement Learning-Based Self-Repair Method for Solving the Agile Satellite Scheduling Problem. Tsinghua Science and Technology, 2026, 31(1): 180-198. https://doi.org/10.26599/TST.2024.9010164
Part of a topical collection:

3538

Views

270

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 07 May 2024
Revised: 10 July 2024
Accepted: 03 September 2024
Published: 25 August 2025
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).