AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (916.4 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Microgrid energy management strategy based on two-stage deep reinforcement learning

Lingxia LUMingjie HUWeiye LUOMiao YU( )
College of Electrical Engineering, Zhejiang University, Hangzhou 310027, China
Show Author Information

Abstract

Objective

With the gradual transformation of the global energy structure and the rapid development of renewable energy technologies, microgrid technology has emerged as an important and rapidly growing area in the energy sector. As the core of microgrid operation and control, the energy management system ensures the efficient use of renewable energy and the stable operation of microgrids through precise monitoring and intelligent control. Traditional energy management methods struggle to effectively handle the complex interrelationships among variables in microgrids, whereas deep reinforcement learning (DRL) enables intelligent decision-making by learning through interaction with the environment and adjusting strategies based on feedback signals. To address the high exploration cost and low training efficiency of existing DRL algorithms, this study proposes a microgrid energy management strategy based on a two-stage DRL framework.

Methods

The proposed strategy includes two stages: offline and online. First, in the offline stage, linear programming is used to obtain the optimal scheduling results of typical days to construct an expert experience library, and imitation learning is then used to pretrain the agents. This stage involves extracting key information from historical data, such as photovoltaic power, wind power, load demand, and electricity prices, and transforming the data into state–action pairs, thereby forming the pretraining foundation for the agents. Subsequently, in the online stage, the agents interact with the real environment to learn the optimal scheduling strategies for nontypical days. During this stage, having accumulated sufficient knowledge in the offline stage, the agents can significantly improve the environmental tracking accuracy and operational economy. A cliff-walk reward mechanism is introduced to ensure that the agents immediately stop exploring after making decisions that violate constraints, thereby reducing the training cost associated with invalid exploration. Concurrently, the proximal policy optimization (PPO) algorithm is introduced to meet the requirements of continuous action spaces and further improve the performance of the agents.

Results

The proposed algorithm has been validated in a typical microgrid system with three PV stations, one wind turbine, one storage system, and flexible loads. The simulation results show that the convergence speed is significantly improved, and the average daily operating cost is reduced by approximately 14% compared with that of single-stage PPO. A comparison with Double Deep Q-Network, Dueling Deep Q-Network, and Distributed Dueling Deep Q-Network further demonstrates the advantages of the proposed method in achieving optimal performance.

Conclusions

In the two-stage DRL, pretraining in the offline stage enables agents to learn general features and strategies, enabling them to adapt more quickly to new environments and tasks in subsequent missions. Training in the online stage helps agents avoid overfitting to task-specific training data, reducing reliance on such data, lowering the risk of overfitting, and ultimately endowing the strategies with better generalization capability and robustness. Overall, compared with single-stage DRL-based microgrid energy management algorithms, the proposed two-stage DRL-based strategy significantly improves the training efficiency and optimal performance of the agents.

CLC number: TM732 Document code: A Article ID: 1002-4956(2026)06-0037-10

References

【1】
【1】
 
 
Experimental Technology and Management
Pages 37-46

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
LU L, HU M, LUO W, et al. Microgrid energy management strategy based on two-stage deep reinforcement learning. Experimental Technology and Management, 2026, 43(6): 37-46. https://doi.org/10.16791/j.cnki.sjg.2026.06.005

6

Views

0

Downloads

0

Crossref

0

Scopus

Received: 14 December 2025
Published: 20 June 2026
© 2026 Experimental Technology and Management. All rights reserved.

This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).