AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Online First

A Multimodal Fusion Framework for Enhanced Exercise Quantification Integrating RFID and Computer Vision

Department of Automation, Tsinghua University, Beijing 100084, China, and also with Information Media Institute, Beijing College of Politics and Law, Beijing 102628, China
School of Software Engineering, Beijing Jiaotong University, Beijing 100044, China
Global Innovation Exchange, Tsinghua University, Beijing 100084, China
College of Intelligence and Computing, Tianjin University, Tianjin 300350, China
Show Author Information

Abstract

The emerging paradigm of embodied intelligence, which emphasizes the tight coupling between physical embodiment and cognitive processes, has opened up new frontiers in Human-Computer Interaction (HCI), particularly for smart personalized exercise monitoring. Despite its significant potential, existing systems often struggle in multi-user environments, where the accuracy and reliability of exercise tracking are severely compromised by the complexity of human behaviors and interactions. To address this critical issue, this paper proposes a multimodal fusion framework that integrates Radio Frequency Identification (RFID) and Computer Vision (CV) technologies for personalized exercise monitoring in multi-user scenarios, which is the first work to enhance interaction-aware perception in such settings. The system workflow consists of three core modules: Data acquisition, perception modeling, and exercise monitoring, which collectively enable comprehensive analysis of individual exercise behaviors. The system leverages multimodal data from RFID tags and a depth camera to jointly identify object ownership and recognize user activities in real time. Extensive experimental results demonstrate that the proposed system achieves an average matching accuracy of 95%, and an average estimation accuracy of 94%.

References

【1】
【1】
 
 
Tsinghua Science and Technology

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Liu Z, Dang F, Liu X, et al. A Multimodal Fusion Framework for Enhanced Exercise Quantification Integrating RFID and Computer Vision. Tsinghua Science and Technology, 2026, https://doi.org/10.26599/TST.2025.9010107

958

Views

55

Downloads

2

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 21 February 2025
Revised: 06 May 2025
Accepted: 17 June 2025
Published: 29 September 2026
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).