AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research | Open Access

MARMOT: masked autoencoder for modeling transient imaging

Siyuan Shen1,2, Ziheng Wang3, Xingyue Peng1 Ruiqian Li1 Qilin Sun3 Shiying Li1 ( )Jingyi Yu1 ( )
School of Information Science and Technology, ShanghaiTech University, Shanghai, China
Lingang Laboratory, Shanghai, China
School of Data Science and School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China

Equal contributors

Show Author Information

Abstract

Pretrained models have recently emerged as effective priors for a wide range of vision and computational imaging problems. Transient measurements have emerged as a new imaging modality, in which a time-resolved sensor records photon counts over time bins. In confocal non-line-of-sight (NLOS) imaging, a hidden scene can only be observed through multi-bounce transient measurements, making reconstruction from sparse scans severely ill-posed. Existing NLOS methods are often task-specific and typically rely on dense measurements or dedicated supervision for each downstream application. We present MARMOT, a masked autoencoder for modeling transient imaging to facilitate NLOS applications in a self-supervised manner. Given a subset of transient histograms sampled on the relay wall, MARMOT encodes the visible measurements with a Transformer encoder and predicts the missing transients with a lightweight Transformer decoder. This design learns a reusable prior over transient measurements while naturally supporting sparse, irregular, and non-uniform scanning patterns. To enable large-scale pretraining, we build TransVerse, a synthetic dataset of one million confocal NLOS transients rendered from 500,000 Objaverse objects. We evaluate MARMOT in two complementary ways. First, the recovered dense transients can be used for transient completion and NLOS reconstruction from sparse scans. Second, the pretrained encoder can be transferred to downstream visual inference tasks, including classification, albedo estimation, and depth estimation. Across synthetic and real-measured datasets, MARMOT achieves competitive performance and strong robustness under high masking ratios, indicating that large-scale self-supervised pretraining provides a practical prior for transient imaging.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 22

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Shen S, Wang Z, Peng X, et al. MARMOT: masked autoencoder for modeling transient imaging. Visual Intelligence, 2026, 4: 22. https://doi.org/10.1007/s44267-026-00125-1

5

Views

0

Crossref

Received: 30 March 2026
Revised: 08 July 2026
Accepted: 09 July 2026
Published: 14 September 2026
© The Author(s) 2026.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.