AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (3.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Large foundation models empowering ummannod aerial vehicle intelligence: Progress, applications and perspectives

Dian SHAO1,2,3( )Chu TANG1,2,3Min CHANG1,2,3Like LIU4Yule WANG1,2,3Hao LI2,5Junqiang BAI1,2,3
Unmanned System Research Institute, Northwestern Polytechnical University, Xi'an 710072, China
National Key Laboratory of Unmanned Aerial Vehicle Technology, Xi'an 710072, China
School of Artificial Intelligence, Northwestern Polytechnical University, Xi'an 710072, China
School of Software, Northwestern Polytechnical University, Xi'an 710072, China
AVIC Chengdu Aircraft Design & Research Institute, Chengdu 610041, China
Show Author Information

Abstract

Large Foundation Models (LFMs), represented by large language models, vision language models, and vision foundation models, are driving a new wave of intelligent evolution for Unmanned Aerial Vehicles (UAVs). Focusing on this trend, the key characteristics and general capabilities of relevant models are first summarized, followed by a categorization of the mainstream embodied architectures driven by them. A comparison is conducted regarding the adaptability and trade-offs of different architectures within the high-dynamic and strongly constrained scenarios of UAVs. Secondly, an analysis is provided on how various LFMs reshape the four core functional elements of UAVs, including perception, planning, control, and interaction, through mechanisms such as open-world understanding, task-level semantic planning, embodied reasoning control, and multi-modal interaction. Furthermore, focusing on high-level cognitive functions driven by LFMs, the mechanisms, implementation pathways, technical limitations, and evaluation paradigms of reasoning, memory, reflection, and imagination in coping with complex UAV scenarios are discussed. The empowerment patterns and frontier progress of LFMs in four typical decision-making tasks are then summarized, including vision-language navigation, active target search, semantic delivery, and swarm intelligent coordination. Finally, core challenges regarding safety risks and protection mechanisms, engineering implementation, and edge deployment are discussed, envisioning future directions in efficient foundation intelligence, the perception-to-cognition transition, and ubiquitous industrial collaboration.

CLC number: V11 Document code: A Article ID: 1000-6893(2026)15-333148-25

References

【1】
【1】
 
 
Acta Aeronautica et Astronautica Sinica

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
SHAO D, TANG C, CHANG M, et al. Large foundation models empowering ummannod aerial vehicle intelligence: Progress, applications and perspectives. Acta Aeronautica et Astronautica Sinica, 2026, 47(15). https://doi.org/10.7527/S1000-6893.2026.33148

1

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 27 November 2025
Revised: 04 January 2026
Accepted: 04 February 2026
Published: 28 February 2026
© 2026 The Journal of Acta Aeronautica et Astronautica Sinica