AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research | Open Access

SAM2-UNet: segment anything 2 makes strong encoder for natural and medical image segmentation

Xinyu Xiong1Zihuang Wu2Shuangyi Tan3Wenxue Li4Feilong Tang5Ying Chen6Siying Li7Jie Ma1Guanbin Li1 ( )
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
School of Computer and Information Engineering, Jiangxi Normal University, Nanchang, China
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Shenzhen, China
Thrust of Robotics and Autonomous Systems, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
AIM Lab, Faculty of IT, Monash University, Melbourne, Australia
School of Future Technology, South China University of Technology, Guangzhou, China
School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen, China
Show Author Information

Abstract

Image segmentation plays an important role in vision understanding. Recently, the emerging vision foundation models continuously achieved superior performance on various tasks. Following such success, in this paper, we prove that the Segment Anything Model 2 (SAM2) can be a strong encoder for U-shaped segmentation models. We propose a simple but effective framework, termed SAM2-UNet, for versatile image segmentation. Specifically, SAM2-UNet adopts the Hiera backbone of SAM2 as the encoder, while the decoder uses the classic U-shaped design. Additionally, adapters are inserted into the encoder to enable parameter-efficient fine-tuning. Preliminary experiments on various downstream tasks, such as camouflaged object detection, salient object detection, marine animal segmentation, mirror detection, and polyp segmentation, demonstrate that our SAM2-UNet can outperform existing specialized state-of-the-art methods with minimal additional complexity.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 2

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Xiong X, Wu Z, Tan S, et al. SAM2-UNet: segment anything 2 makes strong encoder for natural and medical image segmentation. Visual Intelligence, 2026, 4: 2. https://doi.org/10.1007/s44267-025-00106-w

1419

Views

60

Crossref

Received: 14 November 2025
Revised: 20 December 2025
Accepted: 24 December 2025
Published: 28 February 2026
© The Author(s) 2026.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.