AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (13.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Efficient Video Emotion Recognition via Multi-Scale Region-Aware Convolution and Temporal Interaction Sampling

Xiaorui Zhang1,2( )Chunlin Yuan3Wei Sun4Ting Wang5
College of Computer and Information Engineering, Nanjing Tech University, Nanjing, 211816, China
College of Electronic and Information Engineering, Nanjing University of Information Science and Technology, Nanjing, 210044, China
College of Computer Science, Nanjing University of Information Science and Technology, Nanjing, 210044, China
College of Automation, Nanjing University of Information Science and Technology, Nanjing, 210044, China
College of Electrical Engineering and Control Science, Nanjing Tech University, Nanjing, 211816, China
Show Author Information

Abstract

Video emotion recognition is widely used due to its alignment with the temporal characteristics of human emotional expression, but existing models have significant shortcomings. On the one hand, Transformer multi-head self-attention modeling of global temporal dependency has problems of high computational overhead and feature similarity. On the other hand, fixed-size convolution kernels are often used, which have weak perception ability for emotional regions of different scales. Therefore, this paper proposes a video emotion recognition model that combines multi-scale region-aware convolution with temporal interactive sampling. In terms of space, multi-branch large-kernel stripe convolution is used to perceive emotional region features at different scales, and attention weights are generated for each scale feature. In terms of time, multi-layer odd-even down-sampling is performed on the time series, and odd-even sub-sequence interaction is performed to solve the problem of feature similarity, while reducing computational costs due to the linear relationship between sampling and convolution overhead. This paper was tested on CMU-MOSI, CMU-MOSEI, and Hume Reaction. The Acc-2 reached 83.4%, 85.2%, and 81.2%, respectively. The experimental results show that the model can significantly improve the accuracy of emotion recognition.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 1-19

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang X, Yuan C, Sun W, et al. Efficient Video Emotion Recognition via Multi-Scale Region-Aware Convolution and Temporal Interaction Sampling. Computers, Materials & Continua, 2026, 86(2): 1-19. https://doi.org/10.32604/cmc.2025.071043

8

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 30 July 2025
Accepted: 17 October 2025
Published: 09 December 2025
© The Author 2025.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.