AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (29.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Transformer-Driven Multimodal for Human-Object Detection and Recognition for Intelligent Robotic Surveillance

Aman Aman Ullah#,1,2Yanfeng Wu#,1Shaheryar Najam3Nouf Abdullah Almujally4Ahmad Jalal5,6( )Hui Liu1,7,8( )
Guodian Nanjing Automation Co., Ltd., Nanjing, 210003, China
Department of Biomedical Engineering, Riphah International University, I-14, Islamabad, 44000, Pakistan
Department of Electrical Engineering, Bahria University, H-11, Islamabad, 44000, Pakistan
Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, 11671, Saudi Arabia
Department of Computer Science, Air University, E-9, Islamabad, 44000, Pakistan
Department of Computer Science and Engineering, College of Informatics, Korea University, Seoul, 02841, Republic of Korea
Jiangsu Key Laboratory of Intelligent Medical Image Computing, School of Artificial Intelligence (School of Future Technology), Nanjing University of Information Science and Technology, Nanjing, 210003, China
Cognitive Systems Lab, University of Bremen, Bremen, 28359, Germany

#These authors contributed equally to this work

Show Author Information

Abstract

Human object detection and recognition is essential for elderly monitoring and assisted living however, models relying solely on pose or scene context often struggle in cluttered or visually ambiguous settings. To address this, we present SCENET-3D, a transformer-driven multimodal framework that unifies human-centric skeleton features with scene-object semantics for intelligent robotic vision through a three-stage pipeline. In the first stage, scene analysis, rich geometric and texture descriptors are extracted from RGB frames, including surface-normal histograms, angles between neighboring normals, Zernike moments, directional standard deviation, and Gabor-filter responses. In the second stage, scene-object analysis, non-human objects are segmented and represented using local feature descriptors and complementary surface-normal information. In the third stage, human-pose estimation, silhouettes are processed through an enhanced MoveNet to obtain 2D anatomical keypoints, which are fused with depth information and converted into RGB-based point clouds to construct pseudo-3D skeletons. Features from all three stages are fused and fed in a transformer encoder with multi-head attention to resolve visually similar activities. Experiments on UCLA (95.8%), ETRI-Activity3D (89.4%), and CAD-120 (91.2%) demonstrate that combining pseudo-3D skeletons with rich scene-object fusion significantly improves generalizable activity recognition, enabling safer elderly care, natural human–robot interaction, and robust context-aware robotic perception in real-world environments.

References

【1】
【1】
 
 
Computers, Materials & Continua
Article number: 56

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ullah AA, Wu Y, Najam S, et al. Transformer-Driven Multimodal for Human-Object Detection and Recognition for Intelligent Robotic Surveillance. Computers, Materials & Continua, 2026, 87(1): 56. https://doi.org/10.32604/cmc.2025.072508

3

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 28 August 2025
Accepted: 29 October 2025
Published: 10 February 2026
© The Author 2026.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.