Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
A significant challenge remains in identifying and modeling complex 3D point cloud indoor scenes, especially when indoor structural points exhibit similar spatial geometry. By combining local point convolution and global sparse vector attention (SVA), we propose a network to improve the accuracy of semantic segmentation for challenging and often ambiguous wall surfaces and neighboring components with similar structural geometry. Specifically, maximum priori attention (MPA) enhances focus on critical points to optimize feature aggregation in local multiscale multilayer perception (MLP). Then we propose a PCA-based 3D projective point convolution (PPC), capable of orientating the principal direction of the neighborhood and identifying the projection plane. And the neighborhood points are projected adaptively as convolution kernel points, which can be adjusted dynamically to accommodate neighborhood geometric variations and differences. A more flexible convolution whose weights are learned using the projection kernel points is then executed to emphasize the primary structural features and accommodate geometric variations. Globally, SVA effectively models the interdependence and spatial constraint between distant points. The weighted addition vector representation can adaptively adjust the feature channel weight across different dimensions, highlighting critical features while reducing the loss of fine details caused by scalarized compression. By focusing only on relevant points, sparse attention significantly reduces computation burden without sacrificing accuracy. Experimental results show that our proposed model achieves over 73% mean intersection over union (mIoU) on the S3DIS dataset, maintaining excellent recognition performance for ceilings and floors while significantly improving the segmentation of walls and adjacent structures such as doors, windows, boards, and bookcases, which are commonly prone to confusion. These results confirm the model’s robustness and effectiveness in handling complex indoor environments.
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.
Comments on this article