AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

TDFNet: twice decoding Ⅴ-Mamba-CNN Fusion features for building extraction

Wenlong Wanga Peng Yua ( )Mengmeng Lib Xiaojing ZhongcYuanrong HeaHua Sub Yunxuan Zhoud
College of Computer and Information Engineering, Xiamen University of Technology, Xiamen, China
Key Laboratory of Spatial Data Mining and Information Sharing of Ministry of Education, The Academy of Digital China, Fuzhou University, Fuzhou, China
College of Harbour and Coastal Engineering, Jimei University/Xiamen Key Laboratory of Green and Smart Coastal Engineering, Xiamen, China
State Key Laboratory of Estuarine and Coastal Research, East China Normal University, Shanghai, China
Show Author Information

Abstract

Building extraction from remote sensing imagery is vital for various human activities. But it is challenging due to diverse building appearances and complex backgrounds. Research shows the importance of both global context and spatial details for accurate building extraction. Therefore, methods integrating convolutional neural networks (CNNs) and visual transformers (ViTs) are popular nowadays. However, current methods combining these two methods inadequately merge their features and only perform decoding once, leading to issues like unclear boundaries, internal voids, and susceptibility to non-building elements in complex scenarios with low inter-class and high intra-class variability. To address these issues, this paper introduces a novel extraction method called TDFNet. We first replace ViT with Ⅴ-Mamba, which has linear complexity, and combine it with CNN for feature extraction. A bidirectional fusion module (BFM) is then designed to comprehensively integrate spatial details and global information, thereby enabling accurate identification of boundaries between adjacent buildings, and maintaining the structural integrity of buildings to avoid internal holes. During the decoding process, we propose an Encoder-Decoder Fusion Module (EDFM) to initially merge features from different stages of the encoder and decoder, thereby diminishing the model’s susceptibility to non-building elements with features similar to those of buildings, and consequently reducing the incidence of erroneous extractions. Subsequently, a twice decoding strategy is implemented to enhance the learning of multi-scale features significantly, thereby mitigating the impact of tree occlusions and shadows. Our method yields the state-of-the-art (SOTA) performance on three public building datasets.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 19-38

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Wang W, Yu P, Li M, et al. TDFNet: twice decoding Ⅴ-Mamba-CNN Fusion features for building extraction. Geo-Spatial Information Science, 2026, 29(1): 19-38. https://doi.org/10.1080/10095020.2025.2514812

0

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 27 December 2024
Accepted: 28 May 2025
Published: 11 July 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.