AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (90.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Review | Open Access

A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models

Weiqiang Jin#,1,2Yang Liu#,2Yang Gao#,1Shixiang Tang2Yanghao Zhou3Jinhu Qi4Wentao Zhang4Junli Wang5Jing Gao2Yue Ma4Ziwei Zhang1( )Biao Zhao2( )
Institute of SRIICL, Xi’an Jiaotong University, Xi’an, China
School of Information and Communications Engineering, Xi’an Jiaotong University, Xi’an, China
Department of Electrical and Computer Engineering, National University of Singapore, Kent Ridge, Singapore
Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China
School of Computer Science and Technology, University of Science and Technology of China, Hefei, China
Show Author Information

Abstract

Large Vision-Language Models (LVLMs) have achieved strong performance in multimodal perception, understanding, and generation, but their ability to perform complex logical reasoning remains insufficiently understood. In particular, it is still unclear whether current LVLMs can reliably conduct explicit logical operations, multi-step inference, abstract relational reasoning, and cross-modal evidence integration. Reasoning abilities such as deductive, inductive, abductive, multi-hop, and causal inference are fundamental to robust decision making, trustworthy interaction, and real-world deployment, yet they have not been systematically examined in the LVLM literature. Existing surveys mainly discuss mathematical reasoning, general multimodal intelligence, or benchmark progress, but they do not provide a unified account of complex logical reasoning in LVLMs, including its definition, reasoning types, modeling paradigms, evaluation protocols, and unresolved limitations. To address this gap, this survey develops a unified analytical framework for complex logical reasoning in LVLMs. This survey provides a structured review of this emerging area. We first formalize complex logical reasoning in multimodal settings and organize the literature into five recurrent reasoning families: deductive, inductive, abductive, multi-hop, and causal reasoning. We then review reasoning-oriented LVLM architectures, including unified, modular, and tool-augmented paradigms, and summarize major reasoning mechanisms such as chain-of-thought, program-based reasoning, self-correction, and interpretability-oriented analysis. We further examine representative benchmarks and evaluation protocols, with particular attention to the mismatch between final-answer accuracy and genuine reasoning validity. Based on empirical evidence from representative LVLMs and datasets, we identify common capability trends, recurring failure modes, and key open challenges. Our analysis shows that current LVLMs still struggle with reasoning faithfulness, long-horizon inference, cross-modal grounding, hallucination control, and process-aware evaluation. Finally, we outline future directions in reasoning-oriented data construction, model design, training strategies, evaluation methodology, and deployment. Overall, this survey offers a unified conceptual framework and technical roadmap for advancing LVLMs from strong perceptual systems toward reliable multimodal reasoning agents.

References

【1】
【1】
 
 
Computer Modeling in Engineering & Sciences
Article number: 1

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Jin W, Liu Y, Gao Y, et al. A Comprehensive Review of Complex Logical Reasoning in Large Vision-Language Models. Computer Modeling in Engineering & Sciences, 2026, 148(1): 1. https://doi.org/10.32604/cmes.2026.083586

2

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 10 April 2026
Accepted: 15 June 2026
Published: 27 July 2026
© The Author 2026.

This work is licensed under a Creative Commons Attribution 4.0 International License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.