AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

A Cross-modal Retrieval Method for Remote Sensing Images Integrating Multi-scale Features and Location Information

Zhen Dai, Genping Zhao( ), Lianglun Cheng
School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, Guangdong, China
Show Author Information

Abstract

Accurate cross-modal retrieval of remote sensing images and text depends on effective multimodal feature learning and precise alignment of cross-modal information. However, remote sensing images often exhibit significant variations in target scale, making it challenging for feature learning methods to accurately capture positional information and comprehensively represent both image details and overall semantics. To address this issue, a hybrid network combining CNN and Transformer is proposed to enhance the mode’s capability in semantic representation of both fine-grained details and overall structure of remote sensing images. The network incorporates a multi-scale feature optimization module to improve sensitivity to targets of varying scales, and positional information is embedded to enhance the precise alignment of cross-modal data. Additionally, an adaptive pooling mechanism is introduced after feature learning in both modalities to retain critical semantic information, facilitating better cross-modal semantic alignment. Experimental results on the RSICD and RSITMD datasets show that the proposed method outperforms existing mainstream approaches in recall at K (R@K), significantly improving cross-modal retrieval performance.

CLC number: TP391 Document code: A Article ID: 1007–7162(2026)5–36–13

References

【1】
【1】
 
 
Journal of Guangdong University of Technology
Pages 36-48

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Dai Z, Zhao G, Cheng L. A Cross-modal Retrieval Method for Remote Sensing Images Integrating Multi-scale Features and Location Information. Journal of Guangdong University of Technology, 2026, 43(5): 36-48. https://doi.org/10.12052/gdutxb.250022

3

Views

0

Downloads

0

Crossref

Received: 22 January 2025
Accepted: 26 March 2025
Published: 22 May 2025
© 2026 Editorial Office of Journal of Guangdong University of Technology

This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).