AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Just Accepted

Distilling Vision-Language Models for Explainable Vehicle Collision Prediction

Ruici Zhang( )Yuxiang FengJose EscribanoMohammed Quddus

Centre for Transport Engineering and Modelling, Department of Civil and Environmental Engineering, Imperial College London, London SW7 2AZ, UK.

Show Author Information

Abstract

Vision-based collision warning systems are increasingly recognised as a promising countermeasure against traffic collisions. However, their development is constrained by the limited explainability of deep learning-based collision prediction models. Vision-language models (VLMs), with inherent self-explainability, offer a promising solution, yet two critical challenges persist: (1) how to adapt general-purpose VLMs to the task-specific domain of collision prediction, and (2) how to reduce their high inference latency. To tackle these challenges, this paper develops a novel approach that distils VLMs for explainable vehicle collision prediction. Specifically, a two-step chain-of-thought (CoT) prompting strategy was designed to guide VLMs to first analyse the driving scenario and then predict potential collisions, providing explicit rationales. The VLMs were fine-tuned on both pre-collision and normal driving video clips with descriptive annotations. To reduce latency, a feature-based knowledge distillation approach was introduced to transfer knowledge from the fine-tuned VLM to a smaller one through hidden-state supervision. Experimental results demonstrate that fine-tuning improves prediction accuracy by over 20%, while CoT prompting further enhances both performance and explainability. The distilled VLM reduces inference latency by more than 50% with less than a 5% performance degradation, highlighting its potential to advance vision-based collision warning systems and proactive traffic safety.  

References

【1】
【1】
 
 
Communications in Transportation Research

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang R, Feng Y, Escribano J, et al. Distilling Vision-Language Models for Explainable Vehicle Collision Prediction. Communications in Transportation Research, 2026, https://doi.org/10.26599/COMMTR.2026.9640047

250

Views

30

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 04 February 2026
Revised: 21 June 2026
Accepted: 13 August 2026
Available online: 17 August 2026

©The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0,
http://creativecommons.org/licenses/by/4.0/).