AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Efficient Multiagent Policy Optimization Based on Weighted Estimators in Stochastic Cooperative Environments

College of Intelligence and Computing, Tianjin University, Tianjin 300350, China
National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing 210023, China

A preliminary version of the paper was published in the Proceedings of PRICAI 2018.

Show Author Information

Abstract

Multiagent deep reinforcement learning (MA-DRL) has received increasingly wide attention. Most of the existing MA-DRL algorithms, however, are still inefficient when faced with the non-stationarity due to agents changing behavior consistently in stochastic environments. This paper extends the weighted double estimator to multiagent domains and proposes an MA-DRL framework, named Weighted Double Deep Q-Network (WDDQN). By leveraging the weighted double estimator and the deep neural network, WDDQN can not only reduce the bias effectively but also handle scenarios with raw visual inputs. To achieve efficient cooperation in multiagent domains, we introduce a lenient reward network and scheduled replay strategy. Empirical results show that WDDQN outperforms an existing DRL algorithm (double DQN) and an MA-DRL algorithm (lenient Q-learning) regarding the averaged reward and the convergence speed and is more likely to converge to the Pareto-optimal Nash equilibrium in stochastic cooperative environments.

Electronic Supplementary Material

Download File(s)
jcst-35-2-268-Highlights.pdf (1.5 MB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 268-280

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zheng Y, Hao J-Y, Zhang Z-Z, et al. Efficient Multiagent Policy Optimization Based on Weighted Estimators in Stochastic Cooperative Environments. Journal of Computer Science and Technology, 2020, 35(2): 268-280. https://doi.org/10.1007/s11390-020-9967-6

938

Views

11

Crossref

N/A

Web of Science

13

Scopus

0

CSCD

Received: 20 August 2019
Revised: 23 January 2020
Published: 27 March 2020
©Institute of Computing Technology, Chinese Academy of Sciences 2020