Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
This study proposes an offline reinforcement learning (RL) framework based on critic regularized regression (CRR) to optimize speed guidance at signalized intersections under mixed traffic conditions. Using real-world trajectory data combined with signal phase and timing (SPaT) information, the framework learns safe and efficient driving policies without requiring online interactions. A structured data-processing pipeline converts raw vehicle trajectories into Markov decision process (MDP) format, effectively encoding vehicle motion states and signal timing information to facilitate realistic decision-making. Simulation of urban mobility (SUMO) simulation experiments demonstrate substantial improvements over rule-based baselines in safety, comfort, and efficiency: Time-to-collision (TTC) increases from 2.75 to 8.53 s, jerk is reduced by over 50%, and time headway is consistently maintained at approximately 1.68 s. Trajectory visualizations confirm smoother and more adaptive driving behavior. Comparisons with state-of-the-art (SOTA) approaches, including model predictive control (MPC), behavior cloning (BC), twin delayed deep deterministic policy gradient (TD3), and batch constrained Q-learning (BCQ), highlight CRR’s stable performance across multiple evaluation metrics. Ablation studies reveal the critical role of different components in ensuring robust policy behavior, while communication loss simulations demonstrate framework resilience. The method’s computational efficiency, reflected in its considerably shorter runtime than MPC, and its robustness to communication failures reinforce its practicality for real-time deployment.
This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).
Comments on this article