AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (8.6 MB)
Collect
AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Mandarin lip recognition based on MSAF with multimodal task

Yujun RONG1Xianhai WU2( )Fenglin CAI2Tongxin YANG2Penghua LI2,3
China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou 310000, P. R. China
School of Computer Science and Engineering (School of Artificial Intelligence), Chongqing University of Science and Technology, Chongqing 401331, P. R. China
School of Automation, Chongqing University of Posts and Telecommunications, Chongqing 400065, P. R. China
Show Author Information

Abstract

Multimodal lip recognition aims to enhance speech recognition accuracy and robustness by integrating lip movements and speech information, while also aiding specific user groups in communication. However, existing lip-speaking models predominantly focus on English datasets, leaving research on Chinese lip recognition in its nascent stage. Addressing challenges in handling data features across different modalities, integrating these features, and achieving comprehensive fusion of multimodal features, we propose a multimodal split attention fusion audio visual recognition (MSAFVR) model. Through experiments utilizing a Chinese Mandarin lip reading (CMLR) dataset, our model, MSAFVR, demonstrates significant advancements, achieving a remarkable 92.95% accuracy in Chinese lip reading, surpassing state-of-the-art Mandarin lip reading models.

CLC number: TN929 Document code: A Article ID: 1000-582X(2026)04-107-10

References

【1】
【1】
 
 
Journal of Chongqing University
Pages 107-116

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
RONG Y, WU X, CAI F, et al. Mandarin lip recognition based on MSAF with multimodal task. Journal of Chongqing University, 2026, 49(4): 107-116. https://doi.org/10.11835/j.issn.1000-582X.2026.04.010

159

Views

3

Downloads

0

Crossref

0

Scopus

Received: 12 June 2024
Published: 01 April 2026
© Journal of Chongqing University