Highlights
• This model breaks through the cross-modal coding and structured text training technology.
• End-to-end multimodal foundation model training requires only video and structured text.
• The fine-tuning technique based on low-rank matrices and prompt engineering has significantly improved the scene understanding ability.
• This study constructs a video structured standard dataset for multimodal foundation models of road traffic scene understanding.

京公网安备11010802044758号
Comments on this article