Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Generative modeling has emerged as a mainstream paradigm for data synthesis, offering a promising solution to data scarcity and insufficient coverage of long-tail scenarios. Despite remarkable progress in autonomous driving (AD), its potential for traffic data generation from roadside or infrastructure perspectives remains largely underexplored. To this end, we proposed the first traffic-oriented world model termed the traffic world model (TWM) for controllable and high-fidelity multimodal data generation. It was built upon a unified conditional diffusion transformer (cDiT) architecture, which generated a controllable roadside-view image from a structured road-agent layout, and the image was then utilized as the common visual foundation for subsequent multimodal synthesis. TWM jointly supported layout-to-image generation, image-to-video generation, safety-critical event generation, and image-to-light detection and ranging (LiDAR) point-cloud generation, enabling controllable modeling of temporal evolution, traffic participant relationships, and geometric topology in a coherent manner. Extensive experiments on multiple public and proprietary datasets demonstrated that TWM achieves state-of-the-art performance across four traffic scene generation tasks. Relative to the best-performing baseline for each task, TWM reduces FID by 40% for general traffic image generation, Fréchet video distance (FVD) by 42% for traffic video generation, Fréchet inception distance (FID) by 74% for traffic-event image generation, and depth error by 11% for point-cloud generation. The code will be released upon publication.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).
Comments on this article