Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
High-resolution remote sensing (HRS) images often feature objects of varying sizes within the same scene, presenting significant challenges for conventional CNN with fixed-size receptive fields. To address this issue, we propose a multi-scale adaptive learning network (MSALNet) that learns optimal scales in a weakly supervised manner and efficiently fuses multiscale features to enhance feature integration and representation across varying object scales. The MSALNet begins by extracting original features using dilated convolution, effectively capturing information from objects of diverse sizes. It then learns optimal scale parameters from these features to generate scale-transformed representations tailored to different scene contexts. To ensure seamless integration, the original and scale-transformed features are dynamically aligned and fused across multiple network layers. This process produces robust multiscale representations, which are subsequently processed through a fully connected layer with softmax activation for precise scene classification. Extensive experiments on the RSSCN7, AID, and NWPU-RESISC45 datasets demonstrate that MSALNet significantly outperforms traditional CNN, especially in scene categories with pronounced scale variations. These results highlight its robustness and adaptability in addressing complex HRS scenarios.
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.
Comments on this article