Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Infrared small target detection (IRSTD) plays a crucial role in applications such as traffic monitoring systems and maritime rescue. However, existing IRSTD methods face challenges due to their reliance on a single type of data, making them susceptible to noise and deficient in contextual understanding. Additionally, small and limited datasets hinder model generalization and performance in complex scenarios. Previous methods are mostly based on U-Net architectures that are optimized for small-scale data and involve intricate design. These designs often perform well in specific scenarios, but they struggle to generalize effectively in real-world applications. Inspired by leading vision-language models, we propose an MIRSAM (Multimodal Vision-Language Segment Anything Model for Infrared Small Target Detection), the first framework to integrate text modality with image modality for IRSTD in this article. Given the differences in noise and structural information between infrared and natural images, we fine-tune segment anything model (SAM) by designing a contourlet denoising adapter module (CDAM). Integrated into SAM’s image encoder, this module suppresses noise during feature extraction and encoding, enabling efficient adaptation to the infrared domain. To incorporate textual information, we utilize the text encoder of contrastive language-image pre-training (CLIP) to convert text into high-dimensional feature vectors, which then serve as prompts to extract relevant details from the features. In addition, we build the first multimodal IRSTD dataset, IR-TXPair, containing image-text pairs. Experiments on the newly constructed IR-TXPair dataset demonstrate that the proposed MIRSAM outperforms state-of-the-art methods.
This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
Comments on this article