AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Just Accepted

Unstructured Scene Benchmark (USB): Which VLM Performs Better in Autonomous Driving?

Yuan Zheng1,2( )Chenyi Xie1,2Yanrui Chen1,2Wenhao Yu1,2Shen Li3( )Kunsong Shi4Wai Wong5Xu Qu1,2( )Bin Ran2

1 Jiangsu Province Collaborative Innovation Center of Modern Urban Traffic Technologies, Southeast University, Nanjing 211189, China.

2 School of Transportation, Southeast University, Nanjing 211189, China.

3 Department of Traffic Engineering, Tsinghua University, Beijing 100084, China.

4 State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China.

5 Department of Civil and Natural Resources Engineering, University of Canterbury, Christchurch 8041, New Zealand.

Show Author Information

Abstract

Vision-language models (VLMs) offer the potential for unified perception and language-guided decision-making in autonomous driving. However, existing benchmarks predominantly focus on structured road environments and high-quality imagery, leaving limited evidence on model performance (e.g., reasoning, explanation, and decision traceability) under unstructured scenes or degraded sensing conditions. This study develops a trustworthy test and evaluation framework to systematically assess VLM performance in these challenging contexts. Impromptu vision-language-action (VLA) samples are reorganized into six synchronized camera views augmented with vehicle state information, and twenty realistic input perturbations, covering illumination, weather, sensor reliability, and occlusion, are introduced for each scene. Moreover, six original non-open-ended questions are reformulated into traffic decision templates that combine structured choice sets with template-constrained free-text responses. Two types of evaluation methods are formulated. Multiple-choice questions (MCQs) are based on exact answer matching, aggregated with importance weighting according to expert rankings to prioritize planning tasks, and subjective questions (SQs) are graded using tailored, multi-dimensional LLM-based scoring prompts. Experiments are conducted to assesses their performance of seven open-source VLMs, including multiple-choice questions (MCQs) accuracy, reasoning coherence and visual fidelity of their responses to subjective questions (SQs), and their robustness to input perturbations. Overall, the proposed framework addresses a critical evaluation gap for supporting the deployment of VLMs in autonomous driving applications.

Graphical Abstract

References

【1】
【1】
 
 
Communications in Transportation Research

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zheng Y, Xie C, Chen Y, et al. Unstructured Scene Benchmark (USB): Which VLM Performs Better in Autonomous Driving?. Communications in Transportation Research, 2026, https://doi.org/10.26599/COMMTR.2026.9640034

411

Views

47

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 31 December 2025
Revised: 28 April 2026
Accepted: 09 June 2026
Available online: 11 June 2026

© The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0, http://creativecommons.org/licenses/by/4.0/).