AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (897.8 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Leveraging Transfer Learning for Spatio-Temporal Human Activity Recognition from Video Sequences

Umair Muneer Butt1,2( )Hadiqa Aman Ullah2Sukumar Letchmunan1Iqra Tariq2Fadratul Hafinaz Hassan1Tieng Wei Koh3
School of Computer Sciences, Universiti Sains Malaysia, Penang, 1180, Malaysia
Department of Computer Science, The University of Chenab, Gujrat, 50700, Pakistan
Department of Software Engineering and Information System, Universiti Putra Malaysia, Selangor, 43400, Malaysia
Show Author Information

Abstract

Human Activity Recognition (HAR) is an active research area due to its applications in pervasive computing, human-computer interaction, artificial intelligence, health care, and social sciences. Moreover, dynamic environments and anthropometric differences between individuals make it harder to recognize actions. This study focused on human activity in video sequences acquired with an RGB camera because of its vast range of real-world applications. It uses two-stream ConvNet to extract spatial and temporal information and proposes a fine-tuned deep neural network. Moreover, the transfer learning paradigm is adopted to extract varied and fixed frames while reusing object identification information. Six state-of-the-art pre-trained models are exploited to find the best model for spatial feature extraction. For temporal sequence, this study uses dense optical flow following the two-stream ConvNet and Bidirectional Long Short Term Memory (BiLSTM) to capture long-term dependencies. Two state-of-the-art datasets, UCF101 and HMDB51, are used for evaluation purposes. In addition, seven state-of-the-art optimizers are used to fine-tune the proposed network parameters. Furthermore, this study utilizes an ensemble mechanism to aggregate spatial-temporal features using a four-stream Convolutional Neural Network (CNN), where two streams use RGB data. In contrast, the other uses optical flow images. Finally, the proposed ensemble approach using max hard voting outperforms state-of-the-art methods with 96.30% and 90.07% accuracies on the UCF101 and HMDB51 datasets.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 5017-5033

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Butt UM, Ullah HA, Letchmunan S, et al. Leveraging Transfer Learning for Spatio-Temporal Human Activity Recognition from Video Sequences. Computers, Materials & Continua, 2023, 74(3): 5017-5033. https://doi.org/10.32604/cmc.2023.035512

167

Views

15

Downloads

6

Crossref

4

Web of Science

8

Scopus

Received: 24 August 2022
Accepted: 26 October 2022
Published: 31 March 2023
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.