AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (967.4 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Bimodal Iterative Cross-attention Fusion Ensemble Framework

Zhihong Cai1An Zeng1( )Dan Pan2Jiayu Ye1
School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China
School of Electronics and Information Technology, Guangdong Technical Normal University, Guangzhou 510665, China
Show Author Information

Abstract

Alzheimer’s disease (AD) , as a progressive neurodegenerative disorder, presents significant challenges in early diagnosis and clinical intervention. In medical imaging, structural magnetic resonance imaging (sMRI) captures brain atrophy and structural alterations through high-resolution anatomical imaging, while fluorodeoxyglucose positron emission tomography (FDG-PET) effectively reflects functional changes by monitoring cerebral glucose metabolism. These two modalities hold complementary value in detecting AD-related pathological brain changes. However, existing multimodal AD classification models are limited by suboptimal feature fusion, insufficient inter-modal information interaction, and feature distribution discrepancies, hindering their diagnostic utility. To address these issues, a bimodal iterative cross-attention fusion ensemble framework (BICAFEF) is proposed. This framework comprises base classifiers and a meta-classifier. The base classifiers employ ResNet modules to extract features from sMRI and FDG-PET image patches. A spatial feature shrinking (SFS) module, integrating convolutional operations and adaptive aggregation pooling, is designed to reduce inter-modal redundancy and emphasize discriminative features. Additionally, an iterative cross-attention mechanism is constructed to dynamically capture and reinforce global dependencies and complementary information across modalities through multi-round iterations, thereby resolving the challenge of insufficiently exploiting inter-modal synergies and enhancing AD classification performance. To further improve whole-brain classification accuracy, the framework incorporates a meta-classifier to screen and ensemble base classifiers by discarding those with accuracy below 75%, retaining high-performance classifiers to boost robustness and precision. Visualization analyses validate the framework’s focus on critical brain regions, demonstrating its capability to effectively identify AD-related pathological areas in sMRI and PET modalities. Experimental results show that the framework achieves a five-fold classification accuracy (ACC) of 94.3%, sensitivity (SEN) of 92.6%, specificity (SPE) of 96.3%, AUC of 97.5%, and Matthews correlation coefficient (MCC) of 88.7% in AD vs. healthy control (HC) classification, outperforming state-of-the-art multimodal frameworks.

CLC number: TP391 Document code: A Article ID: 1007–7162(2026)2–1–11

References

【1】
【1】
 
 
Journal of Guangdong University of Technology
Pages 1-11

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Cai Z, Zeng A, Pan D, et al. Bimodal Iterative Cross-attention Fusion Ensemble Framework. Journal of Guangdong University of Technology, 2026, 43(2): 1-11. https://doi.org/10.12052/gdutxb.250002

293

Views

1

Downloads

0

Crossref

Received: 05 January 2025
Accepted: 05 March 2025
Published: 17 June 2025
© 2026 Editorial Office of Journal of Guangdong University of Technology

This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).