Pretrained models have recently emerged as effective priors for a wide range of vision and computational imaging problems. Transient measurements have emerged as a new imaging modality, in which a time-resolved sensor records photon counts over time bins. In confocal non-line-of-sight (NLOS) imaging, a hidden scene can only be observed through multi-bounce transient measurements, making reconstruction from sparse scans severely ill-posed. Existing NLOS methods are often task-specific and typically rely on dense measurements or dedicated supervision for each downstream application. We present MARMOT, a masked autoencoder for modeling transient imaging to facilitate NLOS applications in a self-supervised manner. Given a subset of transient histograms sampled on the relay wall, MARMOT encodes the visible measurements with a Transformer encoder and predicts the missing transients with a lightweight Transformer decoder. This design learns a reusable prior over transient measurements while naturally supporting sparse, irregular, and non-uniform scanning patterns. To enable large-scale pretraining, we build TransVerse, a synthetic dataset of one million confocal NLOS transients rendered from 500,000 Objaverse objects. We evaluate MARMOT in two complementary ways. First, the recovered dense transients can be used for transient completion and NLOS reconstruction from sparse scans. Second, the pretrained encoder can be transferred to downstream visual inference tasks, including classification, albedo estimation, and depth estimation. Across synthetic and real-measured datasets, MARMOT achieves competitive performance and strong robustness under high masking ratios, indicating that large-scale self-supervised pretraining provides a practical prior for transient imaging.
- Article type
- Year
- Co-author
Open Access
Research
Issue
Open Access
Research
Issue
Volumetric video is revolutionizing immersive media, with 3D Gaussian Splatting emerging as a key technology due to its unprecedented real-time rendering quality. However, extending this technology to dynamic scenes presents two major challenges: the massive storage and transmission overhead associated with temporal sequences, and a fragmented toolchain ecosystem that hinders efficient research and development. Existing solutions typically focus on isolated stages such as reconstruction or compression, lacking a unified, end-to-end workflow from data acquisition to final viewing. To address these challenges, we propose a comprehensive dynamic Gaussian processing framework that provides a complete, end-to-end pipeline. This framework systematically integrates the entire process, from data acquisition and standardized preprocessing to a suite of diverse dynamic Gaussian reconstruction algorithms. One of its core contributions is a general-purpose compression framework, compatible with the outputs of various reconstruction methods, which significantly reduces the storage footprint of dynamic sequences while maintaining high visual fidelity. To ensure broad accessibility, we have also developed a cross-platform rendering solution that supports high-quality, interactive free-viewpoint experiences on desktop, mobile, and XR devices. Furthermore, to advance the field, we contribute a large-scale, high-quality dynamic human performance capture dataset. Captured with a dense 81-camera array, the dataset comprises over 130 sequences of diverse human motions, including complex interactions with topological changes. Our integrated framework and dataset aim to bridge the entire pipeline from data creation to end-user application, providing a solid foundation for the large-scale adoption and future research of Gaussian Splatting technology.
京公网安备11010802044758号