AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (898.6 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Multi-subspace multimodal sentiment analysis method based on Transformer

Changning TIAN1Yuzheng HE1Di WANG1( )Bo WAN1Xutong GUO2
School of Computer Science and Technology, Xidian University, Xi'an 710071, China
China Electronics Technology Group Corporation 54th Research Institute, Shijiazhuang 050081, China
Show Author Information

Abstract

Multimodal sentiment analysis refers to recognizing the emotions expressed by characters in a video through textual, visual and acoustic information. Most of the existing methods learn multimodal coherence information by designing complex fusion schemes, while ignoring inter-and intra-modal differentiation information, resulting in a lack of information complementary to multimodal fusion representations. To this end, we propose a multi-subspace Transformer fusion network for multimodal sentiment analysis (MSTFN) method. The method maps different modalities to private and shared subspaces to obtain private and shared representations of different modalities, learning differentiated and unified information for each modality. Specifically, the initial feature representations of each modality are first mapped to their respective private and shared subspaces to learn the private representation containing unique information and the shared representation containing unified information in each modality. Second, under the premise of strengthening the roles of textual and audio modalities, a binary collaborative attention cross-modal Transformer module is designed to obtain textual and audio-based tri-modal representations. Then, the final representation of each modality is generated using modal private and shared representations and fused two by two to obtain a bimodal representation to further complement the information of the multimodal fusion representation. Finally, the unimodal representation, bimodal representation, and trimodal representation are stitched together as the final multimodal feature for sentiment prediction. Experimental results on two benchmark multimodal sentiment analysis datasets show that the present method improves on the binary classification accuracy metrics by 0.025 6/0.014 3 and 0.000 7/0.002 3, respectively, compared to the best benchmark method.

CLC number: TP391.1

References

【1】
【1】
 
 
Journal of Northwest University (Natural Science Edition)
Pages 156-167

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
TIAN C, HE Y, WANG D, et al. Multi-subspace multimodal sentiment analysis method based on Transformer. Journal of Northwest University (Natural Science Edition), 2024, 54(2): 156-167. https://doi.org/10.16152/j.cnki.xdxbzr.2024-02-002

559

Views

1

Downloads

0

Crossref

0

CSCD

Received: 09 December 2023
Published: 25 April 2024
© The Editorial Department of Journal of Northwest University(Natural Science Edition)2024.

This is an open access article under the CC BY-NC-ND 4.0 license (https://creativecommons.org/licenses/by-nc-nd/4.0/).