The emergence of multimodal large language models (MLLMs) has laid the foundation for the vision-language-navigation paradigm, which integrates visual perception, natural language understanding, and navigation control within a unified strategic framework. This paradigm has been rapidly adopted in the UAV domain, attempting to enable UAVs to understand natural language instructions, reason in three-dimensional environments, and make flight decisions. Compared with traditional modular navigation approaches, the end-to-end framework based on MLLMs can simultaneously process linguistic and visual signals, directly mapping perceptual information into control commands. However, a systematic review of UAV-VLN remains scarce. This paper presents a comprehensive review of recent advances in this area: from early modular solutions to reason-centric vision-language-action models. It elucidates how visual, linguistic, and control information are progressively integrated to enhance autonomous navigation capabilities. It further summarizes existing datasets and evaluation protocols, including simulation tasks in both indoor and outdoor complex environments as well as real-world UAV flight trajectories, with evaluation metrics covering success rate, time cost, and semantic comprehension. Finally, it identifies key challenges, including difficulties in cross-modal alignment, insufficient real-time responsiveness in dynamic environments, high annotation costs, and poor decision-making robustness in complex scenarios. This paper outlines new pathways and future research directions for UAV autonomous navigation research. It highlights the potential of MLLMs in enhancing intelligent decision-making and interpretability of UAVs, and serves as a reference for research and practice in achieving safe and efficient autonomous flight of UAVs.
- Article type
- Year
- Co-author
With the rapid advancement of new-generation information technologies, the deep integration of smart city infrastructure and intelligent connected vehicles (ICVs) has emerged as a critical engine for the intelligent transformation of urban transportation. Traditional single-vehicle intelligence often faces limitations such as blind spots, occlusions, and limited perception ranges in complex urban scenarios. Therefore, it is essential to move beyond the independent operation of "human-vehicle-road" units to build an integrated, collaborative "vehicle-road-cloud" ecosystem. Such integration is pivotal for resolving information islands, enhancing traffic safety and efficiency through a closed-loop mechanism of "perception-transmission-calculation-control, " and promoting sustainable urban governance by reducing congestion and carbon emissions.
This study systematically reviewed the research progress, technical architectures, and development trends of this "dual-intelligence" synergy. First, from the viewpoints of operational logic and core architecture, the study elucidated how smart cities empower ICVs. By deploying roadside perception networks and mobile edge computing nodes, the infrastructure provides ICVs with beyond-visual-range information and enables high-precision navigation through digital twin technologies and high-definition maps. Conversely, ICVs act as mobile sensors, providing real-time trajectory and status feedback to the cloud, thereby facilitating traffic situational awareness and decision optimization (e.g., adaptive signal control and congestion management). This interaction establishes a robust multi-level vehicle-road-cloud system framework. Second, a comparative analysis of international development paths was performed to reveal distinct strategies. The United States has traditionally prioritized vehicle-side intelligence (e.g., Tesla's vision-based approach) but is increasingly transitioning toward C-V2X communication standards following spectrum reallocation. The European Union focuses on cross-border interoperability and standardization through projects such as C-Roads, emphasizing data privacy under GDPR. Japan and South Korea rely on government-led legislation and integration of automated driving with high-precision 3D mapping (e.g., SIP-adus and K-City). China adopts a "top-down" design with large-scale dual-intelligence pilot zones in cities such as Beijing and Wuhan, promoting the rapid deployment of 5G-V2X and standardized roadside infrastructure. Furthermore, the study deeply explored the key enabling technologies that support this synergy. Integrated Sensing and Communication, which optimizes spectrum and hardware resources through functional fusion (data sharing) and signal fusion (unified waveforms), was analyzed in detail. Moreover, the study analyzed the hierarchical cloud control system, comprising edge, regional, and central clouds. This system balances real-time local control with global data mining and long-term optimization. Additionally, multisensor fusion positioning algorithms were examined to illustrate the integration of GNSS, INS, and LiDAR through loose, tight, or deep coupling mechanisms. Such integration ensures robust centimeter-level positioning, even in GNSS-denied environments like tunnels or urban canyons.
Despite significant achievements, the collaborative development of smart cities and ICVs faces multifaceted challenges. These include technological bottlenecks in automotive-grade chips and algorithm adaptability, barriers in cross-industry protocol compatibility, and prominent risks regarding data security and privacy protection in cross-border transmission. Consequently, future research must focus on several key directions: achieving semantic alignment and unified representation in multimodal sensing fusion to handle heterogeneous data, developing adaptive protocols based on software-defined networks to ensure compatibility, optimizing dynamic edge-cloud computing resource scheduling to meet real-time demands, and constructing active immune network security frameworks to defend against intelligent cyber-attacks. This study provides a comprehensive theoretical reference and technical support for fostering the deep integration and scalable application of smart city infrastructure and ICVs.
Open Access
Review
Issue
Cross-city transfer learning (CCTL) has emerged as a crucial approach for managing the growing complexity of urban data and addressing the challenges posed by rapid urbanization. This paper provides a comprehensive review of recent advances in CCTL, with a focus on its applications in urban computing tasks, including prediction, detection, and deployment. We examine the role of CCTL in facilitating policy adaptation and influencing behavioral change. Specifically, we provide a systematic overview of widely used datasets, including traffic sensor data, GPS trajectory data, online social network data, and map data. Furthermore, we conduct an in-depth analysis of methods and evaluation metrics employed across different CCTL-based urban computing tasks. Finally, we emphasize the potential of cross-city policy transfer in promoting low-carbon and sustainable urban development. This review aims to serve as a reference for future urban development research and promote the practical implementation of CCTLs.
Open Access
Full Length Article
Issue
The Autonomous Modular Bus (AMB) introduces an innovative approach to public transportation by allowing modular buses to dock and undock seamlessly while in motion. This capability effectively alleviates traffic congestion and decreases energy usage through smoother and more efficient vehicle operation. However, achieving autonomous docking for AMBs poses significant challenges, including the need for precise localization in both horizontal and vertical dimensions and the ability to manage dynamic persistent obstacles in close-range scenarios. Existing Light Detection and Ranging (LiDAR)-based Simultaneous Localization and Mapping (SLAM) algorithms, such as LIO-SAM, perform well in static environments but encounter limitations in dynamic scenarios, particularly with occlusions and vertical drift during AMB docking. In this paper, we propose an enhanced LiDAR-Inertial Measurement Unit (IMU) SLAM framework focused on improving localization accuracy and robustness during AMB docking. Key contributions include: (1) A two-stage scan-to-map matching method with ground constraints to reduce z-axis drift; (2) A factor graph optimization strategy integrating IMU roll and pitch constraints and periodic resetting to mitigate long-term drift; (3) A deep learning-based front vehicle detection and point cloud filtering mechanism to reduce occlusion effects. Experimental evaluations on single-vehicle and dual-vehicle datasets demonstrate that our method significantly reduces Absolute Pose Error (APE) and Relative Pose Error (RPE) compared to existing methods. These results highlight the framework's ability to address the unique challenges of AMB docking, therefore helping alleviate traffic congestion and reduce energy consumption.
京公网安备11010802044758号