Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
The emergence of multimodal large language models (MLLMs) has laid the foundation for the vision-language-navigation paradigm, which integrates visual perception, natural language understanding, and navigation control within a unified strategic framework. This paradigm has been rapidly adopted in the UAV domain, attempting to enable UAVs to understand natural language instructions, reason in three-dimensional environments, and make flight decisions. Compared with traditional modular navigation approaches, the end-to-end framework based on MLLMs can simultaneously process linguistic and visual signals, directly mapping perceptual information into control commands. However, a systematic review of UAV-VLN remains scarce. This paper presents a comprehensive review of recent advances in this area: from early modular solutions to reason-centric vision-language-action models. It elucidates how visual, linguistic, and control information are progressively integrated to enhance autonomous navigation capabilities. It further summarizes existing datasets and evaluation protocols, including simulation tasks in both indoor and outdoor complex environments as well as real-world UAV flight trajectories, with evaluation metrics covering success rate, time cost, and semantic comprehension. Finally, it identifies key challenges, including difficulties in cross-modal alignment, insufficient real-time responsiveness in dynamic environments, high annotation costs, and poor decision-making robustness in complex scenarios. This paper outlines new pathways and future research directions for UAV autonomous navigation research. It highlights the potential of MLLMs in enhancing intelligent decision-making and interpretability of UAVs, and serves as a reference for research and practice in achieving safe and efficient autonomous flight of UAVs.
Comments on this article