4D generation has made remarkable pro-gress in synthesizing dynamic 3D objects from input text, images, or video. However, existing methods often represent motion as an implicit deformation field, which limits direct control and editability. To address this, we propose SkeletonGaussian, a novel framework for generating editable, dynamic 3D Gaussians from monocular video input. Our approach introduces a hierarchical, articulated representation that decomposes motion into sparse, rigid motion explicitly driven by a skeleton and fine-grained, non-rigid motion. In detail, we extract a robust skeleton and drive rigid motion via linear blend skinning, followed by hexplane-based refinement for non-rigid deformation, which enhances interpretability and editability. Experimental results show that SkeletonGaussian surpasses existing methods in visual quality while enabling more intuitive articulated motion editing, establishing a new paradigm for controllable 4D generation.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Research Article
Issue
Computational Visual Media 2026, 12(4): 925-939
Published: 22 September 2026
Downloads:1
Total 1
京公网安备11010802044758号