AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (8.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Spatial structure-aware and cross-scale feature modeling network for remote sensing image semantic segmentation

Fangbin Huang( )Yuxuan Guo
School of Computer Science, Nanjing University of Information Science and Technology, Nanjing 210044, China
Show Author Information

Abstract

Remote sensing images exhibit significant spatial geometric characteristics for ground objects such as buildings and roads, while targets within scenes show enormous scale variations, posing challenges to semantic segmentation algorithms' spatial structure modeling capabilities and cross-scale information processing abilities. Traditional methods lack specialized modeling mechanisms for spatial geometric features and suffer from information loss in multi-scale feature fusion. This paper proposes the SC-Net network, addressing these issues through three key technological innovations. First, we designed a feature attention layer where the spatial attention module captures spatial geometric patterns through directional feature decomposition, and the multi-scale attention module preserves feature information at different scales through adaptive pooling strategies. Second, we constructed a three-branch fusion transformer that employs cross-window attention and nine-group feature key-value pair interactions to achieve collaborative modeling of spatial, multi-scale, and global features. Finally, the multi-branch cascaded decoder enhances segmentation boundary accuracy through hierarchical feature fusion strategies. Comprehensive experiments on three standard remote sensing datasets validated the method's superiority. SC-Net achieved 63.04% mean intersection over union (MIOU) on Wuhan dense labeling dataset (WHDLD), 71.57% on Potsdam dataset, and 81.57% on Vaihingen dataset, outperforming state-of-the-art methods such as AerialFormer and SERNet by 0.67–2.12% MIOU. The method particularly demonstrated outstanding performance in scenarios with complex spatial structures and dense multi-scale targets, providing an effective solution for precise remote sensing image interpretation.

References

【1】
【1】
 
 
Electronic Research Archive
Pages 6391-6417

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Huang F, Guo Y. Spatial structure-aware and cross-scale feature modeling network for remote sensing image semantic segmentation. Electronic Research Archive, 2025, 33(10): 6391-6417. https://doi.org/10.3934/era.2025282

233

Views

0

Downloads

2

Crossref

2

Web of Science

0

Scopus

Received: 12 September 2025
Revised: 20 October 2025
Accepted: 21 October 2025
Published: 28 October 2025
©2025 the Author(s), licensee AIMS Press.

This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)