AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (6.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Open Access

Improving Cross-Modal Semantic Alignment with Cross-Modal Joint Semantic Transformer for Multimodal Sentiment Analysis

School of Computer Science and Technology, Hainan University, Haikou 570228, China
School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China
Show Author Information

Abstract

Multimodal Sentiment Analysis (MSA) aims to comprehensively understand human affective states. To achieve this goal, it integrates heterogeneous modalities, including text, audio, and visual information. However, semantic misalignment within multimodal data and the insufficiency of multimodal feature fusion pose challenges to achieving accurate sentiment prediction. To this end, we propose a Cross-modal Joint Semantic Transformer (CJST) model to achieve cross-modal semantic alignment, thereby enhancing sentiment prediction accuracy. First, we design a Singular Value Decomposition (SVD) based cross-modal semantic alignment strategy that can decouple the time and semantic components of unimodal inputs to reduce the impact of misalignment noise and temporal redundancy. Then, a feature-level low-rank multimodal fusion strategy is developed to achieve high-order interactions among semantic features through tensor-based fusion within the low-rank space. Finally, we conduct various experiments on two well-known MSA benchmark datasets. Extensive experimental results indicate that the proposed CJST model outperforms or matches the state-of-the-art methods.

References

【1】
【1】
 
 
Big Data Mining and Analytics
Pages 1341-1353

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang W, Ding J, Liu H, et al. Improving Cross-Modal Semantic Alignment with Cross-Modal Joint Semantic Transformer for Multimodal Sentiment Analysis. Big Data Mining and Analytics, 2026, 9(5): 1341-1353. https://doi.org/10.26599/BDMA.2025.9020110

1166

Views

130

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 06 August 2025
Revised: 24 September 2025
Accepted: 20 October 2025
Published: 20 August 2026
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).