Photorealistic 3D scene generation is challenging due to the scarcity of large-scale, high-quality real-world 3D datasets; manual modeling has complex workflows requiring specialized expertise. These constraints often result in slow iteration cycles, where each modification demands substantial effort, ultimately stifling creativity. We propose a fast, exemplar-driven framework for generating 3D scenes from a single casual input, such as handheld video or drone footage. Our method first leverages 3D Gaussian splatting to robustly reconstruct input scenes with a high-quality 3D appearance model. We then train a per-scene generative cellular automaton to produce a sparse volume of featurized voxels, effectively amortizing scene generation while enabling controllability. A subsequent patch-based remapping step composites the complete scene from the exemplar’s initial 3D Gaussian splats, successfully recovering the appearance statistics of the input scene. The entire pipeline can be trained in less than 10 min for a given exemplar, and generates scenes in 0.5–2 s. Our method enables interactive creation with full user control. We showcase complex 3D generation results produced from real-world exemplars using a self-contained interactive GUI.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Research Article
Issue
Computational Visual Media 2026, 12(4): 907-923
Published: 22 September 2026
Downloads:1
Total 1
京公网安备11010802044758号