AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (13.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Rethinking medical VQA models: Towards data-efficient learning

College of Computer Science and Technology, Harbin Engineering University, Harbin 150001, China
Department of Computing, The Hong Kong Polytechnic University, Hong Kong 999077, China
Show Author Information

Abstract

Medical visual question answering (VQA) is a key medical AI challenge, but scarce data limits progress. Current methods prioritize more pre-training data, overlooking medical data's inherent constraints (ethics, privacy, specialization) which cause slower accumulation of data than public web data. Sole reliance on more data risks rapid performance plateaus. To address these challenges, this paper proposes the cross-modality discriminative pattern identification model (CMDPI), which fundamentally rethinks training methodologies to provide a data-efficient framework for medical VQA. During pre-training, CMDPI identifies inherent biases in conventional techniques under conditions of data scarcity and introduces a co-regularization approach that integrates multiple pretraining techniques for regularization to enhance model generalizability. For fine-tuning, a difference reconstruction mechanism is proposed that effectively preserves unique discriminative features, and head mixup is introduced to further remedy the issue of overfitting. Experimental results demonstrate that CMDPI achieves performance comparable to or surpassing existing methods while requiring substantially less pre-training data. Our work shows the viability of optimizing training paradigms rather than pursuing indiscriminate data scaling for advancing medical VQA systems.

Graphical Abstract

References

【1】
【1】
 
 
Computational Visual Media
Pages 1085-1106

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
He S, Ren D, Pan H, et al. Rethinking medical VQA models: Towards data-efficient learning. Computational Visual Media, 2026, 12(4): 1085-1106. https://doi.org/10.26599/CVM.2026.9450549

15

Views

1

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 08 February 2026
Accepted: 21 April 2026
Published: 22 September 2026
© The Author(s) 2026.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made.

The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.

To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.

To submit a manuscript, please go to https://jcvm.org.