Sort:
Open Access Research Article Issue
TWM: Multimodal data generation engine via traffic world model
Communications in Transportation Research 2026, 6(3): 9640045
Published: 30 September 2026
Abstract PDF (15.5 MB) Collect
Downloads:88

Generative modeling has emerged as a mainstream paradigm for data synthesis, offering a promising solution to data scarcity and insufficient coverage of long-tail scenarios. Despite remarkable progress in autonomous driving (AD), its potential for traffic data generation from roadside or infrastructure perspectives remains largely underexplored. To this end, we proposed the first traffic-oriented world model termed the traffic world model (TWM) for controllable and high-fidelity multimodal data generation. It was built upon a unified conditional diffusion transformer (cDiT) architecture, which generated a controllable roadside-view image from a structured road-agent layout, and the image was then utilized as the common visual foundation for subsequent multimodal synthesis. TWM jointly supported layout-to-image generation, image-to-video generation, safety-critical event generation, and image-to-light detection and ranging (LiDAR) point-cloud generation, enabling controllable modeling of temporal evolution, traffic participant relationships, and geometric topology in a coherent manner. Extensive experiments on multiple public and proprietary datasets demonstrated that TWM achieves state-of-the-art performance across four traffic scene generation tasks. Relative to the best-performing baseline for each task, TWM reduces FID by 40% for general traffic image generation, Fréchet video distance (FVD) by 42% for traffic video generation, Fréchet inception distance (FID) by 74% for traffic-event image generation, and depth error by 11% for point-cloud generation. The code will be released upon publication.

Open Access Issue
3D Environmental Perception Modeling in the Simulated Autonomous-Driving Systems
Complex System Modeling and Simulation 2021, 1(1): 45-54
Published: 30 April 2021
Abstract PDF (18 MB) Collect
Downloads:219

Self-driving vehicles require a number of tests to prevent fatal accidents and ensure their appropriate operation in the physical world. However, conducting vehicle tests on the road is difficult because such tests are expensive and labor intensive. In this study, we used an autonomous-driving simulator, and investigated the three-dimensional environmental perception problem of the simulated system. Using the open-source CARLA simulator, we generated a CarlaSim from unreal traffic scenarios, comprising 15 000 camera-LiDAR (Light Detection and Ranging) samples with annotations and calibration files. Then, we developed Multi-Sensor Fusion Perception (MSFP) model for consuming two-modal data and detecting objects in the scenes. Furthermore, we conducted experiments on the KITTI and CarlaSim datasets; the results demonstrated the effectiveness of our proposed methods in terms of perception accuracy, inference efficiency, and generalization performance. The results of this study will faciliate the future development of autonomous-driving simulated tests.

Total 2