AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (15.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

TWM: Multimodal data generation engine via traffic world model

Zhenyu Zhang1,Z, Jiaqi Wang1,2,3,Z, Chunmian Lin1,2,3 ( ), Lei Yang4, Chuang Zhang5, Zhanwen Liu6, Jianshan Zhou1,2,3, Xuting Duan1,2,3, Kaige Qu1,2,3, Ruifa Luo7, Daxin Tian1,2,3
School of Transportation Science and Engineering, Beihang University, Beijing 102206, China
State Key Lab of Intelligent Transportation System, Beijing 102206, China
Beijing Key Laboratory for Cooperative Vehicle Infrastructure Systems and Safety Control, Beijing 102206, China
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore 639798, Singapore
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing 100084, China
School of Information Engineering, Chang'an University, Xi'an 710000, China
Shenzhen Genvict Technologies Co., Ltd., Shenzhen 518000, China

Zhenyu Zhang and Jiaqi Wang contributed equally to this work.

Show Author Information

Abstract

Generative modeling has emerged as a mainstream paradigm for data synthesis, offering a promising solution to data scarcity and insufficient coverage of long-tail scenarios. Despite remarkable progress in autonomous driving (AD), its potential for traffic data generation from roadside or infrastructure perspectives remains largely underexplored. To this end, we proposed the first traffic-oriented world model termed the traffic world model (TWM) for controllable and high-fidelity multimodal data generation. It was built upon a unified conditional diffusion transformer (cDiT) architecture, which generated a controllable roadside-view image from a structured road-agent layout, and the image was then utilized as the common visual foundation for subsequent multimodal synthesis. TWM jointly supported layout-to-image generation, image-to-video generation, safety-critical event generation, and image-to-light detection and ranging (LiDAR) point-cloud generation, enabling controllable modeling of temporal evolution, traffic participant relationships, and geometric topology in a coherent manner. Extensive experiments on multiple public and proprietary datasets demonstrated that TWM achieves state-of-the-art performance across four traffic scene generation tasks. Relative to the best-performing baseline for each task, TWM reduces FID by 40% for general traffic image generation, Fréchet video distance (FVD) by 42% for traffic video generation, Fréchet inception distance (FID) by 74% for traffic-event image generation, and depth error by 11% for point-cloud generation. The code will be released upon publication.

Graphical Abstract

References

【1】
【1】
 
 
Communications in Transportation Research
Article number: 9640045

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang Z, Wang J, Lin C, et al. TWM: Multimodal data generation engine via traffic world model. Communications in Transportation Research, 2026, 6(3): 9640045. https://doi.org/10.26599/COMMTR.2026.9640045

512

Views

88

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 30 May 2026
Revised: 19 July 2026
Accepted: 05 August 2026
Published: 30 September 2026
© The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).