Abstract
Vision-based collision warning systems are increasingly recognised as a promising countermeasure against traffic collisions. However, their development is constrained by the limited explainability of deep learning-based collision prediction models. Vision-language models (VLMs), with inherent self-explainability, offer a promising solution, yet two critical challenges persist: (1) how to adapt general-purpose VLMs to the task-specific domain of collision prediction, and (2) how to reduce their high inference latency. To tackle these challenges, this paper develops a novel approach that distils VLMs for explainable vehicle collision prediction. Specifically, a two-step chain-of-thought (CoT) prompting strategy was designed to guide VLMs to first analyse the driving scenario and then predict potential collisions, providing explicit rationales. The VLMs were fine-tuned on both pre-collision and normal driving video clips with descriptive annotations. To reduce latency, a feature-based knowledge distillation approach was introduced to transfer knowledge from the fine-tuned VLM to a smaller one through hidden-state supervision. Experimental results demonstrate that fine-tuning improves prediction accuracy by over 20%, while CoT prompting further enhances both performance and explainability. The distilled VLM reduces inference latency by more than 50% with less than a 5% performance degradation, highlighting its potential to advance vision-based collision warning systems and proactive traffic safety.
京公网安备11010802044758号
Comments on this article