AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Knowledge text-guided frequency-spatial multi-scale feature fusion and collaborative relationship learning for UAV tracking

Ning Lia,b Yuanzhi Zhangc Haojun Aid Yuan Raob Mengyun Liub Yun Pengb Xiaoying Wange ( )
School of Artificial intelligence and Big Data, Henan University of Technology, Zhengzhou, China
Institute of Artificial Intelligence, Guangzhou University, Guangzhou, China
Innovation Research Center, Hainan Harbor & Shipping Holding CO., LTD., Haikou, China
School of Cyber Science and Engineering, Wuhan University, Wuhan, China
Information Center, Third Affiliated Hospital of Sun Yat-sen University, Guangzhou, China
Show Author Information

Abstract

Unmanned aerial vehicle (UAV) tracking technology is crucial in resource exploration and military security fields. However, the lack of accurate estimation of the collaborative relationship between the feature changes of moving object and the background hinders accurate tracking in high-altitude scenes. To address this issue, we design a UAV tracking model with knowledge text-guided frequency-spatial multi-scale feature fusion and collaborative relationship learning, called TFST, which can enhance object features and estimate the collaborative relationship between the object and the background in high-altitude scenes with interference such as small objects, object similarity and occlusion, thereby improving the robustness and stability of object tracking. Our TFST tracker mainly consists of a text-image feature alignment and fusion module (TIF), a frequency-spatial domain feature fusion attention module (FSFA), and a local feature mask reconstruction module (LFMR). Specifically, the TIF module introduces text prompts into the image object detection model, selectively fusing image object through text perception to solve cross-modal auxiliary problems and estimate the collaborative relationship between the object and the background. Meanwhile, the proposed LFMR module integrates multi-frequency and multi-scale features to distill spatial features, demonstrating outstanding performance in capturing boundary features and enhancing the tracking precision of small and boundary-blurred object. Finally, we reconstruct the object local feature mask in the LFMR module to enhance the representation of local features within the global collaborative relationship during the training phase, reducing the impact of small objects being occluded. Extensive experiments on four public tracking benchmarks (namely DTB70, UAVDT, VisDrone2018, and UAV123) validate that our tracker achieved 79.5%, 83.8%, and 86.9% tracking performance under object occlusion, object deformation, and similar backgrounds, outperforming related state-of-the-art tracking models. This method can be tested on larger datasets or under different weather conditions, providing a new solution for different application scenarios.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 1755-1773

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Li N, Zhang Y, Ai H, et al. Knowledge text-guided frequency-spatial multi-scale feature fusion and collaborative relationship learning for UAV tracking. Geo-Spatial Information Science, 2026, 29(3): 1755-1773. https://doi.org/10.1080/10095020.2025.2552258

0

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 26 December 2024
Accepted: 20 August 2025
Published: 08 September 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.