The development of unmanned automated vehicles (UAVs) has become a key focus in aerial robotics, fueling the need for navigation systems capable of performing complex and delicate tasks with speed and precision. However, the end-to-end path tracking process often encounters challenges in learning efficiency, generalization, and varying environmental conditions. In this paper, we propose the novel IRL-TP framework for learning-based UAVs’ trajectory planning that employs a deep inverse reinforcement learning (IRL) approach. Firstly, the RL-based path planner must develop a reward function that effectively captures flight safety, collision avoidance, trajectory smoothness, and navigation efficiency within constrained environments filled with numerous obstacles. To achieve optimal results, a deep reward network is constructed to parametrize the unknown reward function, which effectively and implicitly models the satisfaction of multiple objectives. The regularization of entropy through the learned reward function is utilized to optimize the continuous control policy and improve stability and exploration ability during training with a soft actor-critic (SAC) agent. By combining the reward function inference and policy learning processes, the proposed framework empowers the UAVs to mimic expert behavior and create highly generalized navigation strategies in the “potential map”. In experimental environments with a dense obstacle level, our method achieves a success rate of 97.6% while maintaining an instability metric as low as 0.044 throughout the process. Furthermore, the number of episodes needed to converge the parameters was much faster than other methods (~340). The proposed model not only achieves rapid convergence and a reward value 1.6 times higher in the first 200 training episodes and 1.3 times higher after the entire training process, but also demonstrates an impressive inference time of 2.6 ms per step compared to the basic IRL framework. Compared to state-of-the-art methods—including DQN, PPO, SAC, BC, and GAIL—our approach achieves superior trajectory efficiency, enhanced safety margins, smoother motion, and greater training stability, even in complex 3D environments.
- Article type
- Year
- Co-author
Open Access
Article
Issue
Open Access
Article
Issue
The development of Unmanned Surface Vehicles (USVs) has become a key focus in marine robotics, fueling the need for navigation systems capable of performing complex and delicate tasks with speed and precision. However, the end-to-end path tracking process often encounters challenges in learning efficiency, and generalization, and varying environmental conditions. To achieve sample-efficient and robust USV navigation in dynamic maritime environments, the paper proposes a novel hierarchical multi-view latent imitation learning (IL) architecture. By formulating a latent IL objective, the framework disentangles diverse navigation modalities through continuous variables, preventing mode collapse and enhancing behavioral adaptability to non-stationary conditions. High-dimensional multi-view observations are transformed via a ViT-based backbone into compressed latent features to minimize redundant environmental information. These representations are processed by a Mamba-based action encoder, which leverages selective state-space modeling to capture long-term temporal dependencies with high computational efficiency. A UNet-based decoder subsequently forecasts optimal action sequences by synthesizing spatial maps to infer critical environment-agent relationships. This preliminary multi-view latent IL-based trajectory ensures precise tracking and dynamic stability while adhering to physical vehicle constraints. Experimental results validate that this end-to-end approach achieves robust path planning effectiveness, obstacle avoidance capability, and model training efficiency in complex, multi-modal maritime scenarios.
Open Access
Article
Issue
In the realm of unmanned surface vehicle (USV) operations, leveraging environmental factors to enhance situational awareness has garnered significant academic attention. Developing vision systems for USVs presents considerable challenges, mainly due to variable observational conditions and angular vibrations caused by hydrodynamic forces. The paper proposed a novel MDGAN-DIFI network for end-to-end multi-object tracking (MOT), specifically designed for camera systems mounted on USVs. Beyond enhancing traditional MOT models, the proposed MDGAN-DIFI includes preprocessing modules designed to enhance the efficiency of processing input signal quality. Initially, a Deep Iterative Frame Interpolation (DIFI) module is used to stabilize frames in the spatiotemporal domain. Next, an enhanced generative adversarial network (GAN) model is applied to reduce motion blur affecting objects within the field of view. Finally, a YOLO-CSSA architecture combines dual infrared (IR) and RGB data streams to maintain consistent performance across diverse environmental conditions. By synthesizing intermediate frames and restoring blurred details, the framework seeks to stabilize object motion trajectories and recover distinctive appearance features prior to tracking. This approach directly tackles the main causes of tracking failure in maritime environments, such as motion discontinuities and visual degradation. Experimental results demonstrate that the proposed approach outperforms conventional methods in multi-object tracking on USVs, achieving a maximum accuracy (MOTA) of 47.0% and an IDF1 score of 50.1% under challenging operational conditions. Consequently, the proposed multi-object tracking network provides a more robust foundation for subsequent detection and data association processes.
京公网安备11010802044758号