The complex structures of distributed energy systems (DES) and uncertainties arising from renewable energy sources and user load variations pose significant operational challenges. Model predictive control (MPC) and reinforcement learning (RL) are widely used to optimize DES by predicting future outcomes based on the current state. However, MPC’s real-time application is constrained by its computational demands, making it less suitable for complex systems with extended predictive horizons. Meanwhile, RL’s model-free approach leads to suboptimal data utilization, limiting its overall performance. To address these issues, this study proposes an improved reinforcement learning-model predictive control (RL-MPC) algorithm that combines the high-precision local optimization of MPC with the global optimization capability of RL. In this study, we enhance the existing RL-MPC algorithm by increasing the number of optimization steps performed by the MPC component. We evaluated RL, MPC, and the enhanced RL-MPC on a DES comprising a photovoltaic (PV) and battery energy storage system (BESS). The results indicate the following: (1) The twin delayed deep deterministic policy gradient (TD3) algorithm outperforms other RL algorithms in energy cost optimization, but is outperformed in all cases by RL-MPC. (2) For both MPC and RL-MPC, when the mean absolute percentage error (MAPE) of the first-step prediction is 5%, the total cost increases by ~1.2% compared to that when the MAPE is 0%. However, if the accuracy of the initial prediction data remains constant while only the error gradient of the data sequence increases, the total cost remains nearly unchanged, with an increase of only ~0.1%. (3) Within a 12 h predictive horizon, RL-MPC outperforms MPC, suggesting it as a suitable alternative to MPC when high-accuracy prediction data are limited.
- Article type
- Year
- Co-author
During the initial phases of operation following the construction or renovation of existing buildings, the availability of historical power usage data is limited, which leads to lower accuracy in load forecasting and hinders normal usage. Fortunately, by transferring load data from similar buildings, it is possible to enhance forecasting accuracy. However, indiscriminately expanding all source domain data to the target domain is highly likely to result in negative transfer learning. This study explores the feasibility of utilizing similar buildings (source domains) in transfer learning by implementing and comparing two distinct forms of multi-source transfer learning. Firstly, this study focuses on the Higashita area in Kitakyushu City, Japan, as the research object. Four buildings that exhibit the highest similarity to the target buildings within this area were selected for analysis. Next, the two-stage TrAdaBoost.R2 algorithm is used for multi-source transfer learning, and its transfer effect is analyzed. Finally, the application effects of instance-based (IBMTL) and feature-based (FBMTL) multi-source transfer learning are compared, which explained the effect of the source domain data on the forecasting accuracy in different transfer modes. The results show that combining the two-stage TrAdaBoost.R2 algorithm with multi-source data can reduce the CV-RMSE by 7.23% compared to a single-source domain, and the accuracy improvement is significant. At the same time, multi-source transfer learning, which is based on instance, can better supplement the integrity of the target domain data and has a higher forecasting accuracy. Overall, IBMTL tends to retain effective data associations and FBMTL shows higher forecasting stability. The findings of this study, which include the verification of real-life algorithm application and source domain availability, can serve as a theoretical reference for implementing transfer learning in load forecasting.
京公网安备11010802044758号