Accurate medical image segmentation is essential for effective diagnosis and treatment. Previously we proposed PraNet-V1 as a means to enhance polyp segmentation, introducing a reverse attention (RA) module that utilizes background information. However, PraNet-V1 struggles with multi-class segmentation tasks. To address this limitation, we here propose PraNet-V2, which can effectively handle a broader range of tasks, including multi-class segmentation. At the core of PraNet-V2 is our dual-supervised reverse attention (DSRA) module, which incorporates explicit background supervision, independent background modeling, and semantically enriched attention fusion. Our PraNet-V2 framework exhibits strong performance on four polyp segmentation datasets. Moreover, the integration of DSRA into three state-of-the-art semantic segmentation models enables iterative refinement of foreground segmentation, yielding improvements of up to 1.36% in mean Dice score. Jittor code and supplementary materials are available at https://github.com/ai4colonoscopy/PraNet-V2/tree/main/binary_seg/jittor.
- Article type
- Year
- Co-author
Open Access
Short Communication
Issue
Large Foundation Models (LFMs), represented by large language models, vision language models, and vision foundation models, are driving a new wave of intelligent evolution for Unmanned Aerial Vehicles (UAVs). Focusing on this trend, the key characteristics and general capabilities of relevant models are first summarized, followed by a categorization of the mainstream embodied architectures driven by them. A comparison is conducted regarding the adaptability and trade-offs of different architectures within the high-dynamic and strongly constrained scenarios of UAVs. Secondly, an analysis is provided on how various LFMs reshape the four core functional elements of UAVs, including perception, planning, control, and interaction, through mechanisms such as open-world understanding, task-level semantic planning, embodied reasoning control, and multi-modal interaction. Furthermore, focusing on high-level cognitive functions driven by LFMs, the mechanisms, implementation pathways, technical limitations, and evaluation paradigms of reasoning, memory, reflection, and imagination in coping with complex UAV scenarios are discussed. The empowerment patterns and frontier progress of LFMs in four typical decision-making tasks are then summarized, including vision-language navigation, active target search, semantic delivery, and swarm intelligent coordination. Finally, core challenges regarding safety risks and protection mechanisms, engineering implementation, and edge deployment are discussed, envisioning future directions in efficient foundation intelligence, the perception-to-cognition transition, and ubiquitous industrial collaboration.
京公网安备11010802044758号