AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (9.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Human Motion Prediction Based on Multi-Level Spatial and Temporal Cues Learning

Jiayi Geng1Yuxuan Wu1Wenbo Lu2Pengxiang Su1( )Amel Ksibi3Wei Li1Zaffar Ahmed Shaikh4,5Di Gai6
School of Software, Nanchang University, Nanchang, 330000, China
School of Queen Mary, Nanchang University, Nanchang, 330000, China
Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, P.O.Box 84428, Riyadh, 11671, Saudi Arabia
Department of Computer Science and Information Technology, Benazir Bhutto Shaheed University Lyari, Karachi, 75660, Pakistan
School of Engineering, École Polytechnique Fédérale de Lausanne, Lausanne, 1015, Switzerland
School of Mathematics and Computer Science, Nanchang University, Nanchang, 330000, China
Show Author Information

Abstract

Predicting human motion based on historical motion sequences is a fundamental problem in computer vision, which is at the core of many applications. Existing approaches primarily focus on encoding spatial dependencies among human joints while ignoring the temporal cues and the complex relationships across non-consecutive frames. These limitations hinder the model’s ability to generate accurate predictions over longer time horizons and in scenarios with complex motion patterns. To address the above problems, we proposed a novel multi-level spatial and temporal learning model, which consists of a Cross Spatial Dependencies Encoding Module (CSM) and a Dynamic Temporal Connection Encoding Module (DTM). Specifically, the CSM is designed to capture complementary local and global spatial dependent information at both the joint level and the joint pair level. We further present DTM to encode diverse temporal evolution contexts and compress motion features to a deep level, enabling the model to capture both short-term and long-term dependencies efficiently. Extensive experiments conducted on the Human 3.6M and CMU Mocap datasets demonstrate that our model achieves state-of-the-art performance in both short-term and long-term predictions, outperforming existing methods by up to 20.3% in accuracy. Furthermore, ablation studies confirm the significant contributions of the CSM and DTM in enhancing prediction accuracy.

References

【1】
【1】
 
 
Computers, Materials & Continua
Pages 3689-3707

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Geng J, Wu Y, Lu W, et al. Human Motion Prediction Based on Multi-Level Spatial and Temporal Cues Learning. Computers, Materials & Continua, 2025, 85(2): 3689-3707. https://doi.org/10.32604/cmc.2025.066944

477

Views

3

Downloads

1

Crossref

2

Web of Science

2

Scopus

Received: 21 April 2025
Accepted: 24 July 2025
Published: 23 September 2025
© The Author 2024.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.