AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline

An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios

Huaqing Zhang*Hongbin Ma* ( )Xiaofei ZhangLi WangMinglei Han*Hui Chen*Ao Ding*
School of Automation, Beijing Institute of Technology, Beijing 100081, P. R. China
School of Vehicle and Mobility, Tsinghua University, Beijing 100084, P. R. China
School of Mechanical Engineering, Beijing Institute of Technology, Beijing 100081, P. R. China

This paper was recommended for publication in its revised form by editorial board member, Hyo-Sung Ahn.

Show Author Information

Abstract

Aiming at the problem of multi-agent cooperative confrontation in seize-control scenarios, we design an efficient multi-agent policy self-play (EMAP-SP) learning method. First, a multi-agent centralized policy model is constructed to command the agents to perform tasks cooperatively. Considering that the policy being trained and its historical policies usually have poor exploration capability under incomplete information in self-play trainings, the intrinsic reward mechanism based on random network distillation (RND) is introduced in the self-play learning method. In addition, we propose a multi-step on-policy deep reinforcement learning (DRL) algorithm assisted by off-policy policy evaluation (MSOAO) to learn the best response policy in the self-play. Compared with DRL algorithms commonly used in complex decision problems, MSOAO has more efficient policy evaluation capability, and efficient policy evaluation further improves the policy learning capability. The effectiveness of EMAP-SP is fully verified in MiaoSuan wargame simulation system, and the evaluation results show that EMAP-SP can learn the cooperative policy of effectively defeating the Blue side’s knowledge-based policy under incomplete information. Moreover, the evaluations results in DRL benchmark environments also show that the best response policy learning algorithm MSOAO can promote the agent to learn approximately optimal policies.

References

【1】
【1】
 
 
Unmanned Systems
Pages 987-1004

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang H, Ma H, Zhang X, et al. An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios. Unmanned Systems, 2025, 13(4): 987-1004. https://doi.org/10.1142/S230138502550061X

148

Views

0

Crossref

0

Web of Science

1

Scopus

0

CSCD

Received: 03 March 2024
Revised: 12 July 2024
Accepted: 02 August 2024
Published: 30 September 2024
© World Scientific Publishing Company