AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (973.3 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

A multimodal emotion recognition algorithm based on speech, text and facial expression

Xiao WU1Xuan MOU1Yinhua LIU2,3Xiaorui LIU1,2( )
Automation School, Qingdao University, Qingdao 266071, China
Institute of Future, Qingdao University, Qingdao 266071, China
Shandong Key Laboratory of Industrial Control Technology, Qingdao 266071, China
Show Author Information

Abstract

Aiming at the problems of low recognition accuracy and poor generalization ability of current multimodal emotion recognition algorithms in modal feature extraction and information fusion between modalities, a multimodal emotion recognition algorithm based on speech, text and expression is proposed. Firstly, a shallow feature extraction network (Sfen) combined with parallel convolution module (Pconv) is designed to extract the emotional features in speech and text. A modified Inception-ResnetV2 model is adopted to capture the emotional features of expression in video stream. Secondly, in order to strengthen the correlation among modalities, a cross attention module is designed to optimize the fusion between speech and text modalities. Finally, a bidirectional long and short-term memory module based on attention mechanism (BiLSTM-Attention) is used to focus on key information and maintain the temporal correlation between modalities. By comparing the different combinations of the three modalities, it is found that the hierarchical fusion strategy that processes speech and text in advance can obviously improve the accuracy of the model. Experimental results on the public emotion datasets CH-SIMS and CMU-MOSI show that the proposed model achieves higher recognition accuracy than the baseline model, with three-class and two-class accuracy reaching 97.82% and 98.18% respectively, which proves the effectiveness of the model.

CLC number: TP391

References

【1】
【1】
 
 
Journal of Northwest University (Natural Science Edition)
Pages 177-187

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
WU X, MOU X, LIU Y, et al. A multimodal emotion recognition algorithm based on speech, text and facial expression. Journal of Northwest University (Natural Science Edition), 2024, 54(2): 177-187. https://doi.org/10.16152/j.cnki.xdxbzr.2024-02-004

699

Views

7

Downloads

0

Crossref

0

CSCD

Received: 13 October 2023
Published: 25 April 2024
© The Editorial Department of Journal of Northwest University(Natural Science Edition)2024.

This is an open access article under the CC BY-NC-ND 4.0 license (https://creativecommons.org/licenses/by-nc-nd/4.0/).