In endoscopic surgery, the limited field of view and the nonlinear deformation of organs caused by patient movement and respiration significantly complicate the modeling and accurate tracking of soft tissue surfaces from endoscopic image sequences. To address these challenges, we propose a novel Hybrid Triangular Matching (HTM) modeling framework for soft tissue feature tracking. Specifically, HTM constructs a geometric model of the detected blobs on the soft tissue surface by applying the Watershed algorithm for blob detection and integrating the Delaunay triangulation with a newly designed triangle search segmentation algorithm. By leveraging barycentric coordinate theory, HTM rapidly and accurately establishes inter-frame correspondences within the triangulated model, enabling stable feature tracking without explicit markers or extensive training data. Experimental results on endoscopic sequences demonstrate that this model-based tracking approach achieves lower computational complexity, maintains robustness against tissue deformation, and provides a scalable geometric modeling method for real-time soft tissue tracking in surgical computer vision.
- Article type
- Year
- Co-author
Open Access
Article
Issue
Open Access
Article
Issue
Hybrid CNN-Transformer models are widely used in medical image segmentation because they combine CNN-based local feature extraction with Transformer-based global context modeling. Despite their popularity, these models face several challenges, including computational complexity, noise blurring, and information loss. This paper proposes an enhanced convolutional attention network (ECANet) for liver segmentation. ECANet uses a U-shaped architecture with efficient channel-attention-based skip connections. Both the encoder and decoder are constructed using enhanced convolutional Transformer (ECT) blocks, where group convolution is integrated into the convolutional attention module for efficient Token embedding and channel disentanglement, and a Token-wise multi-layer perceptron (MLP) branch is incorporated into the wide-focus module to improve feature representation across channels. Deep supervision and a hybrid of Binary Cross-Entropy (BCE) and Dice loss are used to improve boundary accuracy. We evaluate the proposed model on the publicly available LiTS17 dataset. Experiments show that ECANet outperforms the compared CNN-based and CNN-Transformer baseline models on both quantitative and qualitative measures.
京公网安备11010802044758号