View-guided point cloud completion (ViPC) enhances completion quality by incorporating image modality, yet existing methods often introduce background interference in cross-modal attention, underutilize image cues aligned with missing regions, and employ decoders with limited structural modeling capacity. To address these limitations, in this work, RGK (Reverse to Get the Key), a cross-modal guidance module composed of Reversed Cross Attention (RevCA) and a Missing Region Image-guided Completion Decoder (MRICD) is proposed. RevCA augments standard cross-attention through point-wise dynamic similarity gating and reversed attention redistribution to emphasize features associated with missing regions. MRICD further selects an auxiliary image token sequence based on RevCA and performs cross-attention-based interaction for accurate completion. Experiments on ShapeNet-ViPC show that integrating RGK into baseline networks consistently improves performance over baselines and competing methods, demonstrating its effectiveness in extracting key cross-modal features and restoring point cloud structures.
- Article type
- Year
- Co-author
Open Access
Issue
Open Access
Issue
Existing research has shown that there are hidden features between the global and local features of point cloud, and the representation capability of point cloud can be enhanced by mining and utilizing the hidden features. However, existing theories have not delved deeply into the analysis and utilization of hidden features. To further exploit the hidden features, we propose an Attention-based Hidden Feature Utilization (AHU) module, which consists of two sub-modules. On one hand, a sub-module based on channel attention mechanism enhances the inter-channel dependencies of features, improving the significance of hidden features; on the other hand, another sub-module based on cross-attention mechanism projects the learned hidden features back to the original local features, establishing long-distance dependencies between them and promoting information fusion, such that the generalization ability of the module can be improved. This paper extends the theory about hidden features and the experimental results demonstrate that the AHU module can be integrated into existing state-of-the-art networks to significantly improve the performance.
Open Access
Issue
Transformer tends to take advantage of capturing remote dependencies to extract relational interactions at remote points of the point cloud, ignoring important local structural details, and achieves high performance by significantly increasing the computational burden. To alleviate this problem, we propose a separable Transformer point cloud classification method, named Sep-point, based on the idea of separable visual Transformer. The proposed Sep-point facilitates sequential local-global relational interactions within and between groups of point clouds through depth-separable self-attention. New location token embedding and group self-attention methods are used to compute inter-group attentional relationships with negligible computational cost and to establish telematic interactions across multiple regions, respectively. In this way, the local-global features are extracted while the computational burden is significantly reduced. Experimental results show that the proposed Sep-point improves the classification accuracy by 0.2% on the ModelNet40 dataset over the existing PCT (Point Cloud Transformer) and by 6.3% on the real ScanObjectNN dataset, respectively. Moreover, the number of network parameters and FLOPS metrics are reduced by 0.72M and 0.18G, respectively. These experimental results clearly demonstrate the promising effectiveness of our proposed method.
京公网安备11010802044758号