In the context of the rapidly growing low-altitude economyand the urgent need for high-precision, lightweight UAV object detection in agriculture, logistics, and emergency rescue, this paper tackles the challenges of low target pixel occupancy, environmental occlusion, and severe perspective distortion in low-altitude imagery. Based on the RT-DETR algorithm, an enhanced detection model—Cross-scale Alignment and Position Encoding Enhanced RT-DETR (CAPE-RT-DETR)—is proposed. Firstly, to overcome the limitation of traditional static convolution kernels in terms of feature extraction flexibility under complex backgrounds, this paper proposes a feature enhancement moduleintegrating dynamic convolution kernel generation and gated feature selection, termed C2ML. By utilizing a Large Kernel Predictor (LKP) to dynamically generate spatially adaptive convolution kernels, and combining it with a gated feature selection mechanism to eliminate redundant background information, the module significantly enhances the model’s ability to extract and filter critical features. Secondly, to address the geometric distortion and spatially non-uniform deformation characteristic of aerial perspectives, a learnable positional encoding is integrated with the multi-head self-attention mechanism to construct an enhanced position-aware interaction module, termed AIFP. By learning spatial prior information in an end-to-end manner, this module effectively improves the model’s perceptual sensitivity and localization accuracy with respect to the low-altitude-specific spatial structures. Finally, to resolve the pixel misalignment problem caused by simple upsampling in multi-scale feature fusion, a a cross-scale feature calibration (CSFC) module is introduced. This module utilizes a pyramid scene parsing structure to integrate sparse global context and employs a dual-path convolution and grid sampling mechanism to explicitly compensate for cross-scale alignment biases, thereby achieving consistent representation of semantic information. Experimental results on the ALU and VisDrone2019 datasets demonstrate that CAPE-RT-DETR outperforms the baseline algorithm in terms of parameter count, accuracy, and model size. Meanwhile, ablation experiments validate the effectiveness and synergy of the three improved modules. This research provides a high-precision and lightweight methodological foundation and theoretical support for real-time UAV object detection in complex scenarios.
- Article type
- Year
- Co-author
This paper aims to solve the problem of slow convergence rate and large error of existing intelligent algorithms in the process of optimizing support vector machine to identify risky driving behavior. Firstly, Tent mapping was used to replace the random setting of population initialization of ASO algorithm to increase the diversity and quality of atomic population. Secondly, the hybrid mechanism of dimension-by-dimension pinhole imaging reverse learning and Cauchy mutation was used to improve the diversity of preferred positions of atomic individuals and overcome the problem that ASO algorithm is easy to fall into local optimum and premature convergence. Finally, the adaptive variable spiral search strategy was introduced to improve the atomic individual position update process,so as to improve the global search ability of ASO algorithm, realize the effective balance between global search and local development, and alleviate the problem that ASO algorithm is easy to fall into local optimum and lack of convergence accuracy. Taking the vehicle trajectory data of the exit ramp of Shanghai North Cross Channel as the input, the study used the hybrid strategy to improve the ASO algorithm so as to optimize the LSSVM parameters. And it constructed the classification and identification model of the risk driving behavior of the expressway exit ramp based on IASO-LSSVM. Numerical simulation results show that the average value, standard deviation, best fitness and worst fitness of the numerical simulation results of the IASO algorithm in 12 benchmark test functions are closer to the best optimization value. Compared with ASO-LSSVM and LSSVM, the accuracy, precision, recall and F1 value of risk driving behavior classification and identification results of IASO-LSSVM model increased by 11.5~24.5, 14.1~29.0, 15.1~28.6, 14.7~31.2 percentage points respectively, and the error range was the smallest in different types of risky driving behavior identification results. The accuracy and convergence rate of IASO algorithm are better than those of ASO algorithm, and the IASO-LSSVM model can be used for accurate identification of different types of risk driving behavior, which can provide data support and theoretical basis for reasonable discrimination of vehicle driving trajectory state and formulation of early warning and prevention measures of risk driving behavior.
京公网安备11010802044758号