AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.3 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

A Review of Vision-Language Navigation Models for UAVs: From Perceptual Comprehension to Intelligent Decision-Making

Ziyu WANG1Chenxu DU2Yang LIU3( )
Intelligent Transportation Thrust, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, Guangdong, China
School of Transportation and Logistics, Southwest Jiaotong University, Chengdu 611756, Sichuan, China
School of Vehicle and Mobility, Tsinghua University, Beijing 100084, China
Show Author Information

Abstract

The emergence of multimodal large language models (MLLMs) has laid the foundation for the vision-language-navigation paradigm, which integrates visual perception, natural language understanding, and navigation control within a unified strategic framework. This paradigm has been rapidly adopted in the UAV domain, attempting to enable UAVs to understand natural language instructions, reason in three-dimensional environments, and make flight decisions. Compared with traditional modular navigation approaches, the end-to-end framework based on MLLMs can simultaneously process linguistic and visual signals, directly mapping perceptual information into control commands. However, a systematic review of UAV-VLN remains scarce. This paper presents a comprehensive review of recent advances in this area: from early modular solutions to reason-centric vision-language-action models. It elucidates how visual, linguistic, and control information are progressively integrated to enhance autonomous navigation capabilities. It further summarizes existing datasets and evaluation protocols, including simulation tasks in both indoor and outdoor complex environments as well as real-world UAV flight trajectories, with evaluation metrics covering success rate, time cost, and semantic comprehension. Finally, it identifies key challenges, including difficulties in cross-modal alignment, insufficient real-time responsiveness in dynamic environments, high annotation costs, and poor decision-making robustness in complex scenarios. This paper outlines new pathways and future research directions for UAV autonomous navigation research. It highlights the potential of MLLMs in enhancing intelligent decision-making and interpretability of UAVs, and serves as a reference for research and practice in achieving safe and efficient autonomous flight of UAVs.

CLC number: TP242 Article ID: 1000-565X(2026)06-0173-10

References

【1】
【1】
 
 
Journal of South China University of Technology (Natural Science Edition)
Pages 173-182

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
WANG Z, DU C, LIU Y. A Review of Vision-Language Navigation Models for UAVs: From Perceptual Comprehension to Intelligent Decision-Making. Journal of South China University of Technology (Natural Science Edition), 2026, 54(6): 173-182. https://doi.org/10.12141/j.issn.1000-565X.260003

5

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 04 January 2026
Published: 01 June 2026
© Journal of South China University of Technology(Natural Science Edition)