3D Gaussian splatting demonstrates significant advantages in high-fidelity rendering for offline reconstruction of objects or scenes. However, its integration with existing simultaneous localization and mapping (SLAM) systems, especially in monocular scenarios, still suffers from limited localization accuracy and poor rendering quality. In this paper, we propose a new monocular 3D Gaussian splatting SLAM, which achieves high fidelity online 3D Gaussian splatting based reconstruction given monocular video input with significance-guided pruning. Our key idea is to maintain structural compactness when optimizing the 3D Gaussians, by adaptively pruning them using a global significance evaluation based on multi-dimensional cues such as visibility, opacity, and volume coefficient. Specifically, we use a frame-to-model pipeline that jointly optimizes camera poses and 3D Gaussians within a sliding-window framework, ensuring a globally consistent 3D Gaussian representation for high-fidelity rendering with significance-guided pruning. Furthermore, an online monocular depth estimation model is incorporated to extract depth priors from input images, to effectively initialize the 3D Gaussian attributes for better camera tracking. Extensive experiments on Replica and TUM datasets demonstrate that our approach substantially improves both tracking performance and rendering fidelity, and thus provides state-of-the-art results.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Research Article
Issue
Intrinsic image decomposition decomposes an image into reflectance and shading. It has been applied in image editing, augmented reality, and geometry estimation. However, the complete decoupling between reflectance and shading, as well as the consistency of the reconstructed image with the original image, have become the main challenges in the application of intrinsic image decomposition. To improve the performance of the intrinsic image decomposition algorithm for these two challenges, we propose a novel deep learning framework that works separately to learn features unique to different intrinsic images. Based on this framework, we developed more effective loss functions to strengthen the decoupling of reflectance and shading and to maintain the decomposition without losing as much information of the original image as possible. We trained the network on a mixture of synthetic and real datasets and evaluated the results of the experiments on real datasets. The results show that our proposed method not only outperformed existing state-of-the-art methods in qualitative and quantitative comparisons in terms of reflectance but was also competitive in terms of reconstructed consistency and shading. Finally, we implemented several realistic image-editing applications, and the results were visually superior to other results.
京公网安备11010802044758号