AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (27.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Unstructured scene benchmark (USB): Which VLM performs better in autonomous driving

Yuan Zheng1,2( ), Chenyi Xie1,2, Yanrui Chen1,2, Wenhao Yu1,2, Shen Li3( ), Kunsong Shi4, Wai Wong5, Xu Qu1,2( ), Bin Ran2
Jiangsu Province Collaborative Innovation Center of Modern Urban Traffic Technologies, Southeast University, Nanjing 211189, China
School of Transportation, Southeast University, Nanjing 211189, China
Department of Transportation Engineering, Tsinghua University, Beijing 100084, China
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China
Department of Civil and Environmental Engineering, University of Canterbury, Christchurch 8041, New Zealand
Show Author Information

Abstract

Vision-language models (VLMs) offer the potential for unified perception and language-guided decision-making in autonomous driving. However, existing benchmarks predominantly focus on structured road environments and high-quality imagery, leaving limited evidence on model performance (e.g., reasoning, explanation, and decision traceability) under unstructured scenes or degraded sensing conditions. This study develops a trustworthy test and evaluation framework to systematically assess VLM performance in these challenging contexts. Impromptu vision-language-action (VLA) samples are reorganized into six synchronized camera views augmented with vehicle state information, and twenty realistic input perturbations—covering illumination, weather, sensor reliability, and occlusion—are introduced for each scene. Moreover, six original non-open-ended questions are reformulated into traffic decision templates that combine structured choice sets with template-constrained free-text responses. Two types of evaluation methods are formulated. Multiple-choice questions (MCQs) are based on exact answer matching, aggregated with importance weighting according to expert rankings to prioritize planning tasks, and subjective questions (SQs) are graded using tailored, multidimensional large language model (LLM)-based scoring prompts. Experiments are conducted to assess the performance of seven open-source VLMs, including MCQ accuracy, reasoning coherence, visual fidelity of their responses to SQs, and their robustness to input perturbations. Overall, the proposed framework addresses a critical evaluation gap for supporting the deployment of VLMs in autonomous driving applications.

Graphical Abstract

References

【1】
【1】
 
 
Communications in Transportation Research
Article number: 9640034

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zheng Y, Xie C, Chen Y, et al. Unstructured scene benchmark (USB): Which VLM performs better in autonomous driving. Communications in Transportation Research, 2026, 6(3): 9640034. https://doi.org/10.26599/COMMTR.2026.9640034

809

Views

89

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 31 December 2025
Revised: 28 April 2026
Accepted: 09 June 2026
Published: 30 September 2026
© The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).