AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Home iFuture Article
PDF (13.5 MB)
Collect
AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Original Article | Open Access | Just Accepted

DiffMagicFace: Identity consistent facial editing of real videos

Huanghao Yin1( )Shenkun Xu2Kanle Shi2Junhai Yong1( )Bin Wang1( )

1 Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China

2 Kuaishou Technology, Beijing 100085, China

Show Author Information

Abstract

Text-conditioned image editing has greatly benefit-ted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source video and ensuring consistency of the edited subject across frames. In this paper, we introduce DiffMagicFace, a unique video editing framework that integrates two fine-tuned models for text and image control. These models operate concurrently during inference to produce video frames that maintain identity features while seamlessly aligning with the editing semantics. To ensure the consistency of the edited videos, we develop a dataset comprising images showcasing various facial perspectives for each edited subject. The creation of a data set is achieved through rendering techniques and the subsequent application of optimization algorithms. Remarkably, our approach does not depend on video datasets but still delivers high-quality results in both consistency and content. The excellent effect holds even for complex tasks like talking head videos and distinguishing closely related categories. The videos edited using our framework exhibit parity with videos that are made using traditional rendering software. Through comparative analysis with current state-of-the-art methods, our framework demonstrates superior performance in both visual appeal and quantitative metrics.

References

【1】
【1】
 
 
iFuture

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yin H, Xu S, Shi K, et al. DiffMagicFace: Identity consistent facial editing of real videos. iFuture, 2026, https://doi.org/10.26599/IF.2026.9710003

347

Views

8

Downloads

0

Crossref

Received: 15 April 2026
Revised: 03 July 2026
Accepted: 16 July 2026
Available online: 16 July 2026

© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).