AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Uniformity of markov elements in deep reinforcement learning for traffic signal control

Bao-Lin Ye1,2( )Peng Wu1,2Lingxi Li3Weimin Wu4
School of Information Science and Engineering, Jiaxing University, Jiaxing 314001, China
School of Information Science and Engineering, Zhejiang Sci-Tech University, Hangzhou 310018, China
Elmore Family School of Electrical and Computer Engineering, Purdue University, Indianapolis 46202, USA
State Key Laboratory of Industrial Control Technology, Zhejiang University, Hangzhou 310027, China
Show Author Information

Abstract

Traffic signal control (TSC) plays a crucial role in enhancing traffic capacity. In recent years, researchers have demonstrated improved performance by utilizing deep reinforcement learning (DRL) for optimizing TSC. However, existing DRL frameworks predominantly rely on manually crafted states, actions, and reward designs, which limit direct information exchange between the DRL agent and the environment. To overcome this challenge, we propose a novel design method that maintains consistency among states, actions, and rewards, named uniformity state-action-reward (USAR) method for TSC. The USAR method relies on: 1) Updating the action selection for the next time step using a formula based on the state perceived by the agent at the current time step, thereby encouraging rapid convergence to the optimal strategy from state perception to action; and 2) integrating the state representation with the reward function design, allowing for precise assessment of the efficacy of past action strategies based on the received feedback rewards. The consistency-preserving design method jointly optimizes the TSC strategy through the updates and feedback among the Markov elements. Furthermore, the method proposed in this paper employs a residual block into the DRL model. It introduces an additional pathway between the input and output layers to transfer feature information, thus promoting the flow of information across different network layers. To assess the effectiveness of our approach, we conducted a series of simulation experiments using the simulation of urban mobility. The USAR method, incorporating a residual block, outperformed other methods and exhibited the best performance in several evaluation metrics.

References

【1】
【1】
 
 
Electronic Research Archive
Pages 3843-3866

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ye B-L, Wu P, Li L, et al. Uniformity of markov elements in deep reinforcement learning for traffic signal control. Electronic Research Archive, 2024, 32(6): 3843-3866. https://doi.org/10.3934/era.2024174

0

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 29 February 2024
Revised: 23 April 2024
Accepted: 06 May 2024
Published: 15 June 2024
©2024 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)