AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Single-image 3D Generation via Depth-consistent Supervision and CLIP-based Semantic Alignment

Jian Zhu, Kangxian Huang, Bingfeng Chen, Zhaohui Liao( ), Ruichu Cai, Guang Feng
School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, Guangdong, China
Show Author Information

Abstract

Single-image 3D generation has broad application potential in digital asset creation, virtual reality, and metaverse content production. However, existing optimization methods based on Score Distillation Sampling (SDS) often suffer from geometric ambiguity and multi-view texture inconsistency due to the lack of explicit geometric constraints and semantic supervision. To address these issues, in this research, a two-stage single-image 3D generation method based on 3D Gaussian Splatting (3DGS) is proposed. In the geometry-constrained generation stage, the method combines the SDS framework with multi-view geometric supervision, synthesizing novel-view images via a pre-trained multi-view diffusion model and providing pseudo-depth constraints through a monocular depth estimation network, thereby explicitly enhancing shape reconstruction accuracy and cross-view consistency. In the texture refinement stage, the optimized Gaussian representation is converted into an explicit mesh, and the texture is iteratively refined using a pre-trained diffusion model to recover high-frequency details and local structures. Additionally, a semantic consistency constraint based on the Contrastive Language-Image Pre-training (CLIP) model is introduced to ensure that corresponding semantic regions across multiple views maintain coherent appearance, further improving texture fidelity and visual coherence. Experimental results show that this method significantly improves the accuracy of 3D geometric structures and the quality of texture details while maintaining efficient generation, validating its effectiveness in single-image 3D generation tasks.

CLC number: TP391 Document code: A Article ID: 1007–7162(2026)5–49–8

References

【1】
【1】
 
 
Journal of Guangdong University of Technology
Pages 49-56

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zhu J, Huang K, Chen B, et al. Single-image 3D Generation via Depth-consistent Supervision and CLIP-based Semantic Alignment. Journal of Guangdong University of Technology, 2026, 43(5): 49-56. https://doi.org/10.12052/gdutxb.250191

3

Views

0

Downloads

0

Crossref

Received: 29 October 2025
Accepted: 24 December 2025
Published: 28 February 2026
© 2026 Editorial Office of Journal of Guangdong University of Technology

This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).