AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Research on foundation models for radar remote sensing: progress and prospects

College of Electronic Science and Technology, National University of Defense Technology, Changsha 410073, China
Show Author Information

Abstract

Significance

AI(artificial intelligence) technologies have profoundly transformed the field of remote sensing, revolutionizing data collection, processing, and analysis. Traditionally reliant on manual interpretation and task-specific models, remote sensing has been significantly enhanced by the advent of foundation models. The core of foundational models for radar remote sensing lies in pre-training on massive and multi-source radar remote sensing data to establish a base with universal representation capabilities, thereby achieving breakthroughs in three key aspects. Firstly, it significantly enhances the model's transfer learning and generalization ability across various tasks such as terrain classification, change detection, and deformation inversion, effectively mitigating the constraint of scarce annotated data. Secondly, it supports a unified multi-task modeling framework, which reduces the cost of algorithm development and enhances the profound understanding of complex terrain structures and dynamic processes. Thirdly, it facilitates in-depth integration of multi-modal data including radar and optical imagery, enabling comprehensive all-weather, all-time-phase, and high-precision perception. This technical paradigm not only accelerates the intellectualization and automation of remote sensing information processing, providing strong support for major applications such as disaster emergency response, environmental monitoring, and national defense security, but also promotes open data sharing, model interoperability, and standardization efforts. It is progressively emerging as a core infrastructure for artificial intelligence in remote sensing in the future.

Progress

Recently, researchers have developed a series of foundation models for radar remote sensing interpretation tasks, the construction of which has been supported by radar remote sensing datasets. Current datasets can be categorized into radar remote sensing visual datasets and visual-language datasets based on their modalities. Visual datasets are constructed by unifying and standardizing multiple public datasets to form large-scale multi-category pre-training datasets, such as SARDet-100K, RSAR, and MuSID. Visual-language datasets, on the other hand, utilize multimodal large language models (e.g., ChatGPT-4o) for text information annotation and the generation of question-answer dialogue instructions, such as EarthDial-Instruct, SARLANG-1M, and SARChat-2M. Currently, foundation models applied to radar remote sensing images can be divided into three categories: radar remote sensing visual foundation models, radar remote sensing visual-language foundation models, and physics-informed radar remote sensing foundation models. In the construction of radar remote sensing visual foundation models, methods such as masked image modeling and contrastive learning are employed to effectively uncover the intrinsic structures and patterns within large-scale data, thereby enhancing the capability for generic feature extraction and reducing reliance on annotated data. Representative examples include SARATR-X, CROMA, and AnySat. Radar remote sensing visual-language foundation models aim to unify the representation of image and natural language information. Their core lies in pre-training on large-scale image-text data, enabling the model to accomplish unified multi-task interpretation through textual interaction, thereby improving the ability to handle complex problems in radar remote sensing interpretation. Representative models include SARChat-Bench-2M and SARLANG-1M, which provide methodologies for data collection, annotation, pre-training, and evaluation in the radar image domain where textual annotation is challenging, offering insights for other vertical domains. The information contained in radar remote sensing data reflects the response of ground objects to radar beams. Researchers leverage physical knowledge, such as the physical mechanisms and geometric characteristics of radar remote sensing images (e.g., statistical features, scattering characteristics, and polarimetric domain features), to guide the construction and training of foundation models. Representative methods include FG-MAE, SUMMIT, and RingMoE. In summary, current research methodologies for radar remote sensing foundation models are based on the pre-training and fine-tuning paradigm. At different stages, the geometric and scattering properties of radar remote sensing images are incorporated to improve the model training framework. Through self-supervised learning and constraints informed by physical knowledge, these models are driven to learn generic feature representations of radar remote sensing images.

Conclusions and Prospects

Foundation models with generalized capabilities are of great significance to the development of intelligent interpretation in radar remote sensing. This paper reviews the basic concepts and characteristics of foundation models, the key technologies involved in their construction, and commonly used evaluation methods. Furthermore, it summarizes the current research and application landscapes of vision foundation models for radar remote sensing, vision-language foundation models for radar remote sensing, and physically-informed foundation models for radar remote sensing. On this basis, the challenges in constructing and applying radar remote sensing foundation models are analyzed, and prospects are provided in terms of model architecture, interpretability, lightweight design, and security. In summary, the application of foundation models in the field of radar remote sensing represents a major advancement in intelligent remote sensing interpretation, significantly enhancing the capability for intelligent interpretation and application of remote sensing data, and marking a significant step forward in empowering the field of remote sensing with artificial intelligence.

CLC number: TN957.52 Document code: A Article ID: 1001-2486(2026)02-382-14

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 382-395

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
WANG X, CHEN S. Research on foundation models for radar remote sensing: progress and prospects. Journal of National University of Defense Technology, 2026, 48(2): 382-395. https://doi.org/10.11887/j.issn.1001-2486.25050009

541

Views

9

Downloads

0

Crossref

0

Web of Science

1

Scopus

0

CSCD

Received: 08 May 2025
Published: 01 April 2026
© 2026 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).