Publications
Sort:
Open Access Issue
Attention mechanism and intelligent processing of radar images: progress and prospects
Journal of National University of Defense Technology 2026, 48(3): 36-51
Published: 01 June 2026
Abstract PDF (3.9 MB) Collect
Downloads:1
Significance

SAR is capable of acquiring high-resolution two-dimensional radar images through azimuthal aperture synthesis and range pulse compression, presenting the geometric structure and scattering characteristics of the observed area intuitively. As an active microwave imaging remote sensing device, SAR features strong penetration capability, all-day and all-weather operation, and immunity to illumination and weather conditions, playing an irreplaceable role in military reconnaissance, disaster assessment, marine rights protection, and other critical fields. However, inherent speckle in radar images blurs target features and distorts texture details, while the coupling of target and background scattering properties in complex scenes significantly increases the difficulty of intelligent target recognition. Inspired by the human visual system’s selective attention mechanism, attention mechanisms have been introduced into deep learning-based radar image processing, enabling models to adaptively assign weights to focus on critical information and suppress redundant interference. Given the rapid development and wide application of attention mechanisms in this domain, a systematic review of relevant research progress is of great academic value and practical significance to promote continuous innovation and in-depth engineering application of radar image intelligent processing.

Progress

This paper systematically combed the development context of attention mechanisms, which can be divided into four stages: recurrent attention models, explicit spatial feature selection, channel-wise feature calibration, and self-attention dominated Transformer architectures. Typical attention models were categorized into four types: channel attention, spatial attention, self-attention, and hybrid attention, with their core principles and representative structures elaborated. On this basis, the innovative applications of various attention mechanisms were comprehensively reviewed in key radar image processing tasks, including preprocessing, target detection, image segmentation, target recognition, change detection, multi-modal fusion, and image restoration. To verify the practical performance of attention mechanisms in engineering scenarios, a comparative experiment was conducted on radar image target detection based on YOLOv11s, using HRSID and SAR-AIRcraft-1.0 datasets. Five representative attention mechanisms (GAM, RFAConv, CoT, SCSA, and MLCA) were evaluated in terms of precision, recall, mAP50, mAP50:95, model parameters, computational complexity, and real-time performance. Experimental results show that RFAConv achieves the highest performance gain in ship target detection, while GAM, MLCA, and CoT exhibit respective advantages in precision, recall, and multi-scale localization accuracy for aircraft target detection.

Conclusions and Prospects

In conclusion, attention mechanisms effectively enhance the feature learning ability and task performance of radar image intelligent processing by selectively focusing on critical information and suppressing clutter interference. Future research directions are prospected from four aspects: First, improve the interpretability of attention mechanisms by establishing the mapping relationship between attention weights and radar scattering physical mechanisms to break the black-box limitation of data-driven models. Second, design efficient attention architectures via sparse modeling, dimension decomposition, and physical prior embedding to meet the real-time and lightweight requirements of resource-constrained platforms such as airborne and spaceborne systems. Third, optimize attention mechanisms for multi-modal fusion by constructing dynamic weight allocation and cross-modal semantic correlation modeling to fully exploit complementary information among heterogeneous data sources. Fourth, develop physics-guided attention designs tailored for radar-specific foundation models to address representation bias and semantic gaps caused by direct migration of general vision Transformers, supporting the development of large-scale, high-performance radar remote sensing foundation models.

Open Access Issue
Research on foundation models for radar remote sensing: progress and prospects
Journal of National University of Defense Technology 2026, 48(2): 382-395
Published: 01 April 2026
Abstract PDF (5.2 MB) Collect
Downloads:11
Significance

AI(artificial intelligence) technologies have profoundly transformed the field of remote sensing, revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, remote sensing has been significantly enhanced by the advent of foundation models. The core of foundational models for radar remote sensing lies in pre-training on massive and multi-source radar remote sensing data to establish a base with universal representation capabilities, thereby achieving breakthroughs in three key aspects. Firstly, it significantly enhances the model's transfer learning and generalization ability across various tasks such as terrain classification, change detection, and deformation inversion, effectively mitigating the constraint of scarce annotated data. Secondly, it supports a unified multi-task modeling framework, which reduces the cost of algorithm development and enhances the profound understanding of complex terrain structures and dynamic processes. Thirdly, it facilitates in-depth integration of multi-modal data including radar and optical imagery, enabling comprehensive all-weather, all-time-phase, and high-precision perception. This technical paradigm not only accelerates the intellectualization and automation of remote sensing information processing, providing strong support for major applications such as disaster emergency response, environmental monitoring, and national defense security, but also promotes open data sharing, model interoperability, and standardization efforts. It is progressively emerging as a core infrastructure for artificial intelligence in remote sensing in the future.

Progress

Recently, researchers have developed a series of foundation models for radar remote sensing interpretation tasks, the construction of which has been supported by radar remote sensing datasets. Current datasets can be categorized into radar remote sensing visual datasets and visual-language datasets based on their modalities. Visual datasets are constructed by unifying and standardizing multiple public datasets to form large-scale multi-category pre-training datasets, such as SARDet-100K, RSAR, and MuSID. Visual-language datasets, on the other hand, utilize multimodal large language models (e.g., ChatGPT-4o) for text information annotation and the generation of question-answer dialogue instructions, such as EarthDial-Instruct, SARLANG-1M, and SARChat-2M. Currently, foundation models applied to radar remote sensing images can be divided into three categories: radar remote sensing visual foundation models, radar remote sensing visual-language foundation models, and physics-informed radar remote sensing foundation models. In the construction of radar remote sensing visual foundation models, methods such as masked image modeling and contrastive learning are employed to effectively uncover the intrinsic structures and patterns within large-scale data, thereby enhancing the capability for generic feature extraction and reducing reliance on annotated data. Representative examples include SARATR-X, CROMA, and AnySat. Radar remote sensing visual-language foundation models aim to unify the representation of image and natural language information. Their core lies in pre-training on large-scale image-text data, enabling the model to accomplish unified multi-task interpretation through textual interaction, thereby improving the ability to handle complex problems in radar remote sensing interpretation. Representative models include SARChat-Bench-2M and SARLANG-1M, which provide methodologies for data collection, annotation, pre-training, and evaluation in the radar image domain where textual annotation is challenging, offering insights for other vertical domains. The information contained in radar remote sensing data reflects the response of ground objects to radar beams. Researchers leverage physical knowledge, such as the physical mechanisms and geometric characteristics of radar remote sensing images (e.g., statistical features, scattering characteristics, and polarimetric domain features), to guide the construction and training of foundation models. Representative methods include FG-MAE, SUMMIT, and RingMoE. In summary, current research methodologies for radar remote sensing foundation models are based on the pre-training and fine-tuning paradigm. At different stages, the geometric and scattering properties of radar remote sensing images are incorporated to improve the model training framework. Through self-supervised learning and constraints informed by physical knowledge, these models are driven to learn generic feature representations of radar remote sensing images.

Conclusions and Prospects

Foundation models with generalized capabilities are of great significance to the development of intelligent interpretation in radar remote sensing. This paper reviews the basic concepts and characteristics of foundation models, the key technologies involved in their construction, and commonly used evaluation methods. Furthermore, it summarizes the current research and application landscapes of vision foundation models for radar remote sensing, vision-language foundation models for radar remote sensing, and physically-informed foundation models for radar remote sensing. On this basis, the challenges in constructing and applying radar remote sensing foundation models are analyzed, and prospects are provided in terms of model architecture, interpretability, lightweight design, and security. In summary, the application of foundation models in the field of radar remote sensing represents a major advancement in intelligent remote sensing interpretation, significantly enhancing the capability for intelligent interpretation and application of remote sensing data, and marking a significant step forward in empowering the field of remote sensing with artificial intelligence.

Total 2