AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (11.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

MoTIF: An end-to-end multimodal road traffic scene understanding foundation model

Zihe Wanga,cHaiyang Yua,cChangxin Chena,cZhiyong Cuia,c( )Yufeng Bib,cYilong Rena,cZijian Wangb,cDelan Kongb,cJing TiandShoutong YuaneZhiqiang Lie
School of Transportation Science and Engineering, Beihang University, Beijing, 100083, China
Shandong Hi-speed Group Co., Ltd., Jinan, 250098, China
State Key Laboratory of Intelligent Transportation System, Beijing, 100191, China
School of Transportation Management, People's Public Security University of China, Beijing, 100038, China
CATARC Intelligent and Connected Technology Co., Ltd., Tianjin, 300380, China
Show Author Information

Highlights

• This model breaks through the cross-modal coding and structured text training technology.

• End-to-end multimodal foundation model training requires only video and structured text.

• The fine-tuning technique based on low-rank matrices and prompt engineering has significantly improved the scene understanding ability.

• This study constructs a video structured standard dataset for multimodal foundation models of road traffic scene understanding.

Abstract

Video-based road intelligent detection constitutes a critical component in modern intelligent transportation systems, serving as a crucial role for comprehensive transportation planning and emergency traffic management. Current traffic scene perception methodologies relying on conventional deep learning architectures present inherent limitations, including heavy dependence on extensive manual annotations of specific traffic scenarios and predefined rule configurations. These approaches demonstrate constrained semantic representation capacity and limited generalizability across heterogeneous traffic scenarios. To address these challenges, this study proposes a novel end-to-end multimodal foundation model architecture that jointly generates dynamic traffic event detection outcomes and semantic-rich contextual descriptions. Through integration of low-rank adaptation (LoRA) and prompt fine-tuning as parameter-efficient fine-tuning strategies, we develop the multimodal road traffic scene understanding foundation model (MoTIF), which establishes cross-modal alignment between visual patterns and textual semantics. This framework demonstrates enhanced capability in extracting salient traffic targets and generating hierarchical scene representations, significantly improving automated detection efficiency in road video analytics. Notably, MoTIF exhibits contextual reasoning capabilities for implicit traffic event interpretation. Extensive evaluations on two real-world datasets encompassing urban road intersection scenarios in Tianjin and highway monitoring systems in Shandong Province reveal that MoTIF achieves superior performance metrics: 65.81 average score on multimodal scene understanding assessment and 83.33% event detection accuracy, outperforming mainstream benchmarks in both precision and computational efficiency. This research advances multimodal learning paradigms for intelligent transportation systems while providing practical insights for adaptive traffic management applications.

Graphical Abstract

References

【1】
【1】
 
 
Communications in Transportation Research
Article number: 100227

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Wang Z, Yu H, Chen C, et al. MoTIF: An end-to-end multimodal road traffic scene understanding foundation model. Communications in Transportation Research, 2025, 5(4): 100227. https://doi.org/10.1016/j.commtr.2025.100227

1506

Views

105

Downloads

3

Crossref

4

Web of Science

6

Scopus

Received: 02 June 2025
Revised: 21 August 2025
Accepted: 24 August 2025
Published: 08 December 2025
© 2025 The Authors.

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).