Publications
Sort:
Issue
Multiagent reinforcement learning-based coordinated control for hybrid transformers
Journal of Tsinghua University (Science and Technology) 2026, 66(8): 1715-1725
Published: 31 August 2026
Abstract PDF (2.1 MB) Collect
Downloads:0
Objective

Flexible interconnection technology can realize asynchronous closed-loop operation among multiple alternating current distribution feeders. It can support dedicated power flow transfer, balance feeder loading, reduce network losses, and enable optimal allocation of controllable resources. The thyristor-controlled hybrid transformer (TCHT) can be used to realize flexible interconnection in distribution networks, offering good economic performance and high reliability. In device operation and control, the power-transfer command—whether obtained from an optimal power flow algorithm or specified manually based on engineering experience—must be converted into voltage compensation tap positions that are executable by the TCHT. However, in practical engineering, the accurate determination of the equivalent circuit parameters of lines, transformers, and other components is challenging. Existing studies have shown that a feedback mechanism can be introduced for a single TCHT, whereby the compensation voltage that minimizes power transfer error can be determined through a shortest-path search method. However, this method is difficult to extend to coordinate control when multiple TCHTs are coupled through the network topology.

Methods

This study proposed a coordinated operation and control method based on multiagent reinforcement learning (RL). First, the coordinated control of multiple TCHTs was modeled as a Markov decision process (MDP). The system state space was defined as the local power regulation error of each TCHT. The action space was defined as the neighboring points of the current three-phase voltage compensation tap position of the TCHT, which avoids the convergence complexities of the training and control processes caused by an excessively large action space. The reward function was defined as the maximum regulation error. Consequently, the original problem was transformed into an equivalent MDP. Second, an online solution method based on multiagent RL was developed. A parameterized policy function was adopted to handle the continuous state space. Generalized advantage estimation was applied to reduce the approximation variance of the policy gradient, after which policy parameters were updated via the proximal policy optimization algorithm. Based on this, a coordinated policy training algorithm for multiple TCHTs was developed. Guided by the global reward function, this method enabled coordination among different TCHT devices.

Results

Numerical analyses were performed on typical public network topologies. A flexible simulation platform was constructed for the interconnection distribution network based on Python and pandapower, with three homogeneous TCHT devices installed in the process. The results showed that the policy iteration process of the proposed method was stable, with only small fluctuations. The reward value increased significantly from iterations 0 to 50. From iterations 50 to 350, the value converged gradually to the optimum. From iterations 350 to 400, it remained basically stable around the optimum. Additionally, each decision step required only one forward pass of the neural network. On an Intel i7-13700 processor, each computation required only 10 ms on average, meeting real-time requirements. Compared with the independent shortest-path search method, the proposed method reduced the average regulation steps and error by 36.6% and 58.9%, respectively.

Conclusions

These results show that although existing methods cannot coordinate multiple TCHTs within a common distribution network in the absence of power grid parameters, our algorithm significantly improves the accuracy and efficiency of coordinated control. Thus, this study fills the methodological gap in the collaborative control problem of multiple TCHTs.

Total 1