AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research | Open Access

Benchmarking Drag★ for eye direction transformation and beyond

Yuxiang Fu1Jiakun Ding2Renzhi Wang3Qian Fu4Ivor W. Tsang5Ming-Ming Cheng1Qing Guo1,6 ( )
VCIP, College of Computer Science, Nankai University, Tianjin, China
College of Intelligence and Computing, Tianjin University, Tianjin, China
Department of Electrical and Computer Engineering, University of Alberta, Edmonton, Canada
Data61, CSIRO, Canberra, Australia
IHPC and CFAR, Agency for Science, Technology, and Research, Singapore, Singapore
College of Design and Engineering, National University of Singapore, Singapore, Singapore
Show Author Information

Abstract

Eye direction plays a crucial role in determining the quality of photographs containing human faces. Images where subjects look in inconsistent directions are often perceived as low-quality and discarded. While state-of-the-art deep generative models such as DragGAN, DragDiffusion, and DragonDiffusion (collectively referred to as Drag★) offer potential solutions for eye direction transformation, their effectiveness for this specific task remains unexplored. In this work, we systematically investigate the capability of Drag ★ for eye direction transformation. Our initial experiments reveal that these models in their original form cannot effectively perform this task. To address this limitation, we construct a specialized dataset (i.e., eye multi-direction dataset (EMDD)) and establish a comprehensive benchmark for evaluating methods through fine-tuning on our curated data. Our analysis demonstrates that fine-tuned models achieve satisfactory results when the angular difference between the directions of the source and target eyes is small. However, we observe significant performance degradation when large directional changes are necessary. Through detailed investigation, we uncover the underlying causes of these limitations and provide insights into the models’ failure modes. To overcome these challenges, we propose the edge-localized point selector and zero-latent source region replacement, which can alleviate the identified limitations. Experimental results demonstrate that our approach achieves substantial performance improvements for eye direction transformation, particularly in scenarios involving large angular changes.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 29

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Fu Y, Ding J, Wang R, et al. Benchmarking Drag★ for eye direction transformation and beyond. Visual Intelligence, 2025, 3: 29. https://doi.org/10.1007/s44267-025-00103-z

375

Views

2

Crossref

Received: 01 September 2025
Revised: 01 December 2025
Accepted: 02 December 2025
Published: 28 February 2026
© The Author(s) 2025.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.