AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (12.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Causal Cross-Modal Context Fusion for Real-Time Video Summarization with Predictive Tracking and Validated Adaptive Evaluation

Aravapalli Rama Satish1Sai Babu Veesam2( )Shonak Bansal3( )Krishna Prakash4Mohammad Rashed Iqbal Faruque5( )
School of Computer Science and Engineering, VIT-AP University, Amaravati, India
Department of AI & DS, KLEF Deemed to be University, Vaddeswaram, Guntur, India
Department of Electronics and Communication Engineering, Chandigarh University, Gharuan, Mohali, India
Department of Information and Communication Technology, Marwadi University, Rajkot, India
Space Science Centre (ANGKASA), Institute of Climate Change (IPI), Universiti Kebangsaan Malaysia, Bangi, Malaysia
Show Author Information

Abstract

Real-time video streams now flood everything from security cameras to social media, yet current summarization systems still stumble when audio, visual, and semantic cues unfold with tangled cause–and–effect patterns. Most cross-modal transformers treat correlations as if time were a flat canvas, ignoring how an early sound might trigger a later visual event in the process. They also lack mechanisms to predict tracking uncertainty, adapt to narrative shifts, or evolve their own evaluation criteria, leaving summaries brittle and often incoherent in process. To address these gaps, we propose a Cross-Modal Context Fusion framework built from five tightly linked components. A Temporal-Causal Graph Memory Network captures directional cause–and–effect edges across audio, video, and semantic signals, improving the logical flow of detected key segments. Additionally, a Predictive Entropy Reinforcement Engine learns camera focus and keyframe decisions that minimize future uncertainty, thereby stabilizing tracking under rapid motion or noise. The Cross-Modal Residual Synergy Transformer explicitly models discrepancies, such as off-screen speech, and feeds those residuals back to refine the fusion. For long-form narrative coherence, a Dynamic Hierarchical Context Predictor alternates between micro-actions and macro-story arcs, balancing fine detail with global structure. Finally, a Self-Evolving Evaluation Loop meta-learns to adjust loss weights as deployment contexts shift, sustaining performance without costly full retraining sets. Experiments on SumMe, TVSum, and long-form documentaries indicate up to 15% F1 and 12% ROUGE-L gains, with human studies reporting 18% higher perceived coherence and >90% sustained approval in process. The result is a video summarizer that reasons causally, anticipates uncertainty, adapts its own metrics, and delivers concise yet narratively faithful summaries suited for demanding real-time applications.

References

【1】
【1】
 
 
Computer Modeling in Engineering & Sciences
Article number: 35

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Satish AR, Veesam SB, Bansal S, et al. Causal Cross-Modal Context Fusion for Real-Time Video Summarization with Predictive Tracking and Validated Adaptive Evaluation. Computer Modeling in Engineering & Sciences, 2026, 147(3): 35. https://doi.org/10.32604/cmes.2026.080403

11

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 09 February 2026
Accepted: 14 May 2026
Published: 30 June 2026
© The Author 2026.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.