AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Video behavior recognition network using multi time-scale convolution

Xijiang CHEN1( )Quanen LIANG1Xianquan HAN2Qing AN3
School of Safety Science and Emergency Management, Wuhan University of Technology, Wuhan 430070, China
Changjiang River Scientific Research Institute, Wuhan 430010, China
School of Artificial Intelligence, Wuchang University of Technology, Wuhan 430223, China
Show Author Information

Abstract

The behavior recognition network based on 2D convolutional usually integrates classification results of multiple video frames to recognize different behaviors, but it can′t extract space-time feature using the 2D convolution kernels. To solve this problem, MTSC (multi time-scale convolution) was proposed based on TSM (temporal shift module), which contained convolution kernels of different scales to fuse the space-time feature from different time scales. By controlling the position that inserting MTSC into ResNet50 network and the parameter setting of MTSC, the optimal behavior recognition network based on MTSC was discussed. Using the PyTorch training model, an experimental study was conducted on a large open source dataset, Something-Something v2. The results show that the behavior recognition network based on MTSC achieves 59.47% Top-1 accuracy, and outperform TSM and other behavior recognition networks.

CLC number: TP391.4 Document code: A Article ID: 1001-2486(2023)03-136-10

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 136-145

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
CHEN X, LIANG Q, HAN X, et al. Video behavior recognition network using multi time-scale convolution. Journal of National University of Defense Technology, 2023, 45(3): 136-145. https://doi.org/10.11887/j.cn.202303016

217

Views

0

Downloads

0

Crossref

0

Web of Science

2

Scopus

1

CSCD

Received: 01 June 2021
Published: 28 June 2023
© 2023 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).