AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Adaptive multi-scale feature aggregation transformer network for single remote sensing image super-resolution

Zhiqi Zhanga,b,c Qi Suna,b Zhiwei Yea,b ( )Chuang Liua,b,c Mi Wangc 
School of Computer Science, Hubei University of Technology, Wuhan, China
Hubei Provincial Key Laboratory of Green Intelligent Computing Power Network, Hubei University of Technology, Wuhan, China
State Key Laboratory of Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan, China
Show Author Information

Abstract

Remote sensing image super-resolution (RSISR) plays a key role in recovering spatial details and improving image quality from satellite imagery. In recent years, transformer-based methods have shown excellent performance in RSISR tasks. Despite the higher computational efficiency of local self-attention calculations compared to global self-attention calculations, its limited receptive field restricts the model from effectively modeling the complex scale diversity and long-range dependencies of ground observation targets. Moreover, the intermediate features of existing methods contain blocking artifacts, leading to different degrees of feature edge distortion and texture detail loss. To address the above issues, this paper proposes the adaptive multi-scale feature aggregation transformer network (AMFAT), which improves the feature representation capability through dynamic weighting and cross-window interaction. Specifically, the adaptive context channel attention (ACCA) is designed to fuse multi-branch features using dynamic weights for object-guided context adaptation. In addition, the mixed-scale token attention (MSTA) is constructed to eliminate blocking artifacts through cross-window interaction. Meanwhile, simple gating units with spatial enhancement operations are introduced into the feed-forward network (FFN) to optimize local feature aggregation. We conducted extensive experiments on four publicly available remote sensing datasets, and the results show that, compared to other methods, AMFAT exhibits excellent performance and adaptability both in terms of quantitative metrics and visual quality. The model will be available at https://github.com/sq-3768/AMFAT.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 1611-1632

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhang Z, Sun Q, Ye Z, et al. Adaptive multi-scale feature aggregation transformer network for single remote sensing image super-resolution. Geo-Spatial Information Science, 2026, 29(3): 1611-1632. https://doi.org/10.1080/10095020.2025.2582378

1

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 14 June 2025
Accepted: 24 October 2025
Published: 11 December 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.