AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Multi-level representation learning via ConvNeXt-based network for unaligned cross-view matching

Fangli Guana Nan Zhaoa Zhixiang FangbLing Jiangc,d,eJianhui Zhanga Yue Yuf ( )Haosheng Huangg 
School of Computer Science, Hangzhou Dianzi University, Hangzhou, China
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan, China
Anhui Province Key Laboratory of Physical Geographic Environment, Chuzhou University, Chuzhou, China
Anhui Engineering Laboratory of Geo-information Smart Sensing and Services, Chuzhou, China
Anhui Center for Collaborative Innovation in Geographical Information Integration and Application, Chuzhou, China
Department of Land Surveying and Geo-Informatics, The Hong Kong Polytechnic University, Hong Kong, China
Department of Geography, Ghent University, Ghent, Belgium
Show Author Information
An erratum to this article is available online at:

Abstract

Cross-view matching refers to the use of images from different platforms (e.g. drone and satellite views) to retrieve the most relevant images, where the key is that the viewpoints and spatial resolution. However, most of the existing methods focus on extracting fine-grained features and ignore the connection of contextual information in the image. Therefore, we propose a novel ConvNeXt-based multi-level representation learning model for the solution of this task. First, we extract global features through the ConvNeXt model. In order to obtain a joint part-based representation learning from the global features, we then replicated the obtained global features, operating one copy with spatial attention and the other copy using a standard convolutional operation. In addition, the features of different branches are aggregated through the multilevel feature fusion module to prepare for cross-view matching. Finally, we created a new hybrid loss function to better limit these features and assist in mining crucial data regarding global features. The experimental results indicate that we have achieved advanced performance on two common datasets, University-1652 and SUES-200 at 89.79% and 95.75% in drone target matching and 94.87% and 98.80 in drone navigation.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 2344-2357

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Guan F, Zhao N, Fang Z, et al. Multi-level representation learning via ConvNeXt-based network for unaligned cross-view matching. Geo-Spatial Information Science, 2025, 28(5): 2344-2357. https://doi.org/10.1080/10095020.2024.2439385

522

Views

9

Crossref

11

Web of Science

11

Scopus

0

CSCD

Received: 07 June 2024
Accepted: 03 December 2024
Published: 17 January 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.