Large Vision-Language Models (LVLMs) have achieved strong performance in multimodal perception, understanding, and generation, but their ability to perform complex logical reasoning remains insufficiently understood. In particular, it is still unclear whether current LVLMs can reliably conduct explicit logical operations, multi-step inference, abstract relational reasoning, and cross-modal evidence integration. Reasoning abilities such as deductive, inductive, abductive, multi-hop, and causal inference are fundamental to robust decision making, trustworthy interaction, and real-world deployment, yet they have not been systematically examined in the LVLM literature. Existing surveys mainly discuss mathematical reasoning, general multimodal intelligence, or benchmark progress, but they do not provide a unified account of complex logical reasoning in LVLMs, including its definition, reasoning types, modeling paradigms, evaluation protocols, and unresolved limitations. To address this gap, this survey develops a unified analytical framework for complex logical reasoning in LVLMs. This survey provides a structured review of this emerging area. We first formalize complex logical reasoning in multimodal settings and organize the literature into five recurrent reasoning families: deductive, inductive, abductive, multi-hop, and causal reasoning. We then review reasoning-oriented LVLM architectures, including unified, modular, and tool-augmented paradigms, and summarize major reasoning mechanisms such as chain-of-thought, program-based reasoning, self-correction, and interpretability-oriented analysis. We further examine representative benchmarks and evaluation protocols, with particular attention to the mismatch between final-answer accuracy and genuine reasoning validity. Based on empirical evidence from representative LVLMs and datasets, we identify common capability trends, recurring failure modes, and key open challenges. Our analysis shows that current LVLMs still struggle with reasoning faithfulness, long-horizon inference, cross-modal grounding, hallucination control, and process-aware evaluation. Finally, we outline future directions in reasoning-oriented data construction, model design, training strategies, evaluation methodology, and deployment. Overall, this survey offers a unified conceptual framework and technical roadmap for advancing LVLMs from strong perceptual systems toward reliable multimodal reasoning agents.
- Article type
- Year
- Co-author
Open Access
Review
Issue
Open Access
Review
Issue
With the growing adoption of Artifical Intelligence (AI), AI-driven autonomous techniques and automation systems have seen widespread applications, become pivotal in enhancing operational efficiency and task automation across various aspects of human living. Over the past decade, AI-driven automation has advanced from simple rule-based systems to sophisticated multi-agent hybrid architectures. These technologies not only increase productivity but also enable more scalable and adaptable solutions, proving particularly beneficial in industries such as healthcare, finance, and customer service. However, the absence of a unified review for categorization, benchmarking, and ethical risk assessment hinders the AI-driven automation progress. To bridge this gap, in this survey, we present a comprehensive taxonomy of AI-driven automation methods and analyze recent advancements. We present a comparative analysis of performance metrics between production environments and industrial applications, along with an examination of cutting-edge developments. Specifically, we present a comparative analysis of the performance across various aspects in different industries, offering valuable insights for researchers to select the most suitable approaches for specific applications. Additionally, we also review multiple existing mainstream AI-driven automation applications in detail, highlighting their strengths and limitations. Finally, we outline open research challenges and suggest future directions to address the challenges of AI adoption while maximizing its potential in real-world AI-driven automation applications.
Open Access
Article
Issue
The UAV pursuit-evasion problem focuses on the efficient tracking and capture of evading targets using unmanned aerial vehicles (UAVs), which is pivotal in public safety applications, particularly in scenarios involving intrusion monitoring and interception. To address the challenges of data acquisition, real-world deployment, and the limited intelligence of existing algorithms in UAV pursuit-evasion tasks, we propose an innovative swarm intelligence-based UAV pursuit-evasion control framework, namely “Boids Model-based DRL Approach for Pursuit and Escape” (Boids-PE), which synergizes the strengths of swarm intelligence from bio-inspired algorithms and deep reinforcement learning (DRL). The Boids model, which simulates collective behavior through three fundamental rules, separation, alignment, and cohesion, is adopted in our work. By integrating Boids model with the Apollonian Circles algorithm, significant improvements are achieved in capturing UAVs against simple evasion strategies. To further enhance decision-making precision, we incorporate a DRL algorithm to facilitate more accurate strategic planning. We also leverage self-play training to continuously optimize the performance of pursuit UAVs. During experimental evaluation, we meticulously designed both one-on-one and multi-to-one pursuit-evasion scenarios, customizing the state space, action space, and reward function models for each scenario. Extensive simulations, supported by the PyBullet physics engine, validate the effectiveness of our proposed method. The overall results demonstrate that Boids-PE significantly enhance the efficiency and reliability of UAV pursuit-evasion tasks, providing a practical and robust solution for the real-world application of UAV pursuit-evasion missions.
京公网安备11010802044758号