Publications
Article type
Sort:
Open Access Article Issue
Improving the segmentation of confusable indoor structures using point convolution and sparse vector attention
Geo-Spatial Information Science 2026, 29(4): 3007-3028
Published: 03 December 2025
Abstract Collect

A significant challenge remains in identifying and modeling complex 3D point cloud indoor scenes, especially when indoor structural points exhibit similar spatial geometry. By combining local point convolution and global sparse vector attention (SVA), we propose a network to improve the accuracy of semantic segmentation for challenging and often ambiguous wall surfaces and neighboring components with similar structural geometry. Specifically, maximum priori attention (MPA) enhances focus on critical points to optimize feature aggregation in local multiscale multilayer perception (MLP). Then we propose a PCA-based 3D projective point convolution (PPC), capable of orientating the principal direction of the neighborhood and identifying the projection plane. And the neighborhood points are projected adaptively as convolution kernel points, which can be adjusted dynamically to accommodate neighborhood geometric variations and differences. A more flexible convolution whose weights are learned using the projection kernel points is then executed to emphasize the primary structural features and accommodate geometric variations. Globally, SVA effectively models the interdependence and spatial constraint between distant points. The weighted addition vector representation can adaptively adjust the feature channel weight across different dimensions, highlighting critical features while reducing the loss of fine details caused by scalarized compression. By focusing only on relevant points, sparse attention significantly reduces computation burden without sacrificing accuracy. Experimental results show that our proposed model achieves over 73% mean intersection over union (mIoU) on the S3DIS dataset, maintaining excellent recognition performance for ceilings and floors while significantly improving the segmentation of walls and adjacent structures such as doors, windows, boards, and bookcases, which are commonly prone to confusion. These results confirm the model’s robustness and effectiveness in handling complex indoor environments.

Total 1