AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (9.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Found-RL: Foundation model-enhanced reinforcement learning via asynchronous VLM feedback for autonomous driving

Yansong Qu1, Zihao Sheng2, Zilin Huang2, Jiancong Chen1, Yuhao Luo2, Tianyi Wang3, Yiheng Feng1, Samuel Labi1, Sikai Chen2( )
Lyles School of Civil and Construction Engineering, Purdue University, West Lafayette 47907, USA
Department of Civil and Environmental Engineering, University of Wisconsin–Madison, Madison 53706, USA
Department of Civil, Architectural and Environmental Engineering, University of Texas at Austin, Austin 78712, USA
Show Author Information

Abstract

Reinforcement learning (RL) has emerged as a dominant paradigm for end-to-end autonomous driving (AD) with real-time inference. However, RL typically suffers from sample inefficiency and a lack of semantic interpretability in complex scenarios. To mitigate these limitations, foundation models (particularly vision-language models (VLMs)) can be integrated because they offer rich, context-aware knowledge. However, deploying such computationally intensive models within high-frequency multienvironment RL training loops is severely hindered by prohibitive inference latency and the absence of unified integration platforms. To bridge this gap, we present Found-RL, a specialized platform tailored to leverage foundation models to efficiently enhance RL for AD. A core innovation of the proposed platform is its asynchronous batch inference framework, which decouples heavy VLM reasoning from the simulation loop. This design effectively resolves latency bottlenecks, supporting real-time or near-real-time RL learning from VLM feedback. Using the proposed platform, we introduce diverse supervision mechanisms to address domain-specific challenges: We first implement value-margin regularization (VMR) and advantage-weighted action guidance (AWAG) to effectively distill expert-like VLM action suggestions into the RL policy. Furthermore, for dense supervision, we adopt high-throughput contrastive language-image pre-training (CLIP) for reward shaping. We mitigate CLIP’s dynamic blindness and probability dilution via conditional contrastive action alignment, which prompts discretized speed/command and yields a normalized, margin-based bonus from context-specific action-anchor scoring. Found-RL delivers an end-to-end pipeline for fine-tuned VLM integration with modular support and shows that a lightweight RL model with millions of parameters can achieve near-VLM performance compared with billion-parameter VLMs while sustaining real-time inference (~500 FPS).

References

【1】
【1】
 
 
Communications in Transportation Research
Article number: 9640027

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Qu Y, Sheng Z, Huang Z, et al. Found-RL: Foundation model-enhanced reinforcement learning via asynchronous VLM feedback for autonomous driving. Communications in Transportation Research, 2026, 6(3): 9640027. https://doi.org/10.26599/COMMTR.2026.9640027

1298

Views

108

Downloads

1

Crossref

0

Web of Science

0

Scopus

Received: 15 February 2026
Revised: 05 April 2026
Accepted: 12 May 2026
Published: 30 September 2026
© The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).