Publications
Article type
Sort:
Open Access Article Issue
Knowledge text-guided frequency-spatial multi-scale feature fusion and collaborative relationship learning for UAV tracking
Geo-Spatial Information Science 2026, 29(3): 1755-1773
Published: 08 September 2025
Abstract Collect

Unmanned aerial vehicle (UAV) tracking technology is crucial in resource exploration and military security fields. However, the lack of accurate estimation of the collaborative relationship between the feature changes of moving object and the background hinders accurate tracking in high-altitude scenes. To address this issue, we design a UAV tracking model with knowledge text-guided frequency-spatial multi-scale feature fusion and collaborative relationship learning, called TFST, which can enhance object features and estimate the collaborative relationship between the object and the background in high-altitude scenes with interference such as small objects, object similarity and occlusion, thereby improving the robustness and stability of object tracking. Our TFST tracker mainly consists of a text-image feature alignment and fusion module (TIF), a frequency-spatial domain feature fusion attention module (FSFA), and a local feature mask reconstruction module (LFMR). Specifically, the TIF module introduces text prompts into the image object detection model, selectively fusing image object through text perception to solve cross-modal auxiliary problems and estimate the collaborative relationship between the object and the background. Meanwhile, the proposed LFMR module integrates multi-frequency and multi-scale features to distill spatial features, demonstrating outstanding performance in capturing boundary features and enhancing the tracking precision of small and boundary-blurred object. Finally, we reconstruct the object local feature mask in the LFMR module to enhance the representation of local features within the global collaborative relationship during the training phase, reducing the impact of small objects being occluded. Extensive experiments on four public tracking benchmarks (namely DTB70, UAVDT, VisDrone2018, and UAV123) validate that our tracker achieved 79.5%, 83.8%, and 86.9% tracking performance under object occlusion, object deformation, and similar backgrounds, outperforming related state-of-the-art tracking models. This method can be tested on larger datasets or under different weather conditions, providing a new solution for different application scenarios.

Total 1