Computed tomography (CT) is currently the most common diagnostic method for brain hemorrhages. Rapidly identifying the location and shape of a brain hemorrhage using deep learning models is of great clinical significance for locating its area and determining its cause. However, most of the mainstream medical segmentation models encounter under-segmentation problems during the segmentation of brain hemorrhages, especially in the region near the skull or when the hemorrhage volume is small. For this reason, this paper proposes a brain hemorrhage segmentation method based on multimodal text representation, which uses a contrastive language-image pretraining (CLIP) model to encode the designed prompts for representation. The text encoder in CLIP can represent information in prompts about the relative location, inclusion relation and other parameters. The representation can then be combined with the U-net to perform the brain hemorrhage segmentation task. The method proposed in this paper uses flexible text prompts to address the problem that some parts of a brain hemorrhage are difficult to segment, thereby enhancing the precision of segmentation. The medical segmentation performance metrics (Dice coefficient) of our method reached 43.3% and 58.8% respectively, when using the publicly available Brain Hemorrhage Segmentation Dataset (BHSD) and the brain hemorrhage dataset for an individual hospital. The improved performance of our method compared with other single medical segmentation models provides strong evidence of its effectiveness.
- Article type
- Year
Open Access
Issue
Open Access
Issue
Segmentation of lumbar vertebrae in CT images has significant importance in the auxiliary diagnosis and treatment of lumbar diseases. The U-Net architecture and its extended models have attracted extensive attention in the field of medical image segmentation. By addressing issues such as the loss of fine-grained features in lumbar segmentation using the Rolling U-Net model, an improved Rolling U-Net model is proposed for segmenting lumbar CT images. This model integrates a convolutional neural network (CNN) with a multi-layer perceptron (MLP). Feature excitation modules are inserted at the fourth convolutional layer and the bottleneck layer to increase the weight of key anatomical structures. By constructing long-range-local blocks (Lo2 blocks), it achieves the fusion of local feature information with long-range dependencies. The core R-MLP module within the Lo2 block learns long-range dependencies across the entire image in a single direction. By controlling and combining R-MLP modules oriented in different directions, OR-MLP and DOR-MLP modules are constructed to capture long-range dependencies in multiple directions. Finally, residual convolutions are integrated to restore segmentation details in the lumbar spine. Simultaneously, a MultiClassDiceCE loss function is designed by combining the pixel classification advantages of the Dice loss function and the cross-entropy loss function. Experimental results indicate that the number of categories and sampling strategies significantly impact the segmentation performance of the improved Rolling U-Net model. For binary segmentation tasks, the division strategy based on the total number of images is recommended to balance accuracy and stability, whereas the division strategy based on the total number of instances is better suited for multi-classification tasks. When performing multi-class segmentation tasks on the JST_LV and VerSe datasets, the improved Rolling U-Net model outperformed segmentation models such as U-Net, Attention U-Net, and Rolling U-Net in terms of average Intersection over Union (IoU), Dice coefficient, recall, specificity, and precision. This demonstrates that the improved model effectively enhances the accuracy, integrity of detail, and classification robustness of lumbar spine CT image segmentation.
京公网安备11010802044758号