AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research | Open Access

SHIELD: an evaluation benchmark for face spoofing and forgery detection with multimodal large language models

Yichen Shi1,Yuhao Gao2,Yingxin Lai3,Hongyang Wang2Jun Feng2Lei He4Jun Wan5Changsheng Chen6Zitong Yu7 ( )Xiaochun Cao8
School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, Shanghai, 200030, China
School of Information Science and Technology, Shijiazhuang Tiedao University, Shijiazhuang, 050043, China
Department of Artificial Intelligence, Xiamen University, Xiamen, 361005, Fujian, China
Electrical Engineering Department, UCLA, Los Angeles, CA 90095, United States of America
New Laboratory of Pattern Recognition(NLPR), Institute of Automation, Chinese Academy of Sciences, Beijing, 100190, China
College of Electronics and Information Engineering, Shenzhen University, Shenzhen, 518060, China
School of Computing and Information Technology, Great Bay University, Dongguan, Guangdong, 523000, China
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University, Shenzhen, 518107, China

Equal contributors

Show Author Information

Abstract

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing and forgery clues in face attack detection tasks remains underexplored. In this paper, we introduce a benchmark, SHIELD, to evaluate MLLMs for face spoofing and forgery detection. Specifically, we design true/false and multiple-choice questions to assess MLLM performance on multimodal face data across two tasks. For the face anti-spoofing task, we evaluate three modalities (i.e., RGB, infrared, and depth) under six attack types. For the face forgery detection task, we evaluate GAN-based and diffusion-based data, incorporating visual and acoustic modalities. We conduct zero-shot and few-shot evaluations in standard and chain of thought (COT) settings. Additionally, we propose a novel multi-attribute chain of thought (MA-COT) paradigm for describing and judging various task-specific and task-irrelevant attributes of face images. The findings of this study demonstrate that MLLMs exhibit strong potential for addressing the challenges associated with the security of facial recognition technology applications.

References

【1】
【1】
 
 
Visual Intelligence
Article number: 9

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Shi Y, Gao Y, Lai Y, et al. SHIELD: an evaluation benchmark for face spoofing and forgery detection with multimodal large language models. Visual Intelligence, 2025, 3: 9. https://doi.org/10.1007/s44267-025-00079-w

1506

Views

23

Crossref

Received: 30 August 2024
Revised: 22 April 2025
Accepted: 23 April 2025
Published: 03 June 2025
© The Author(s) 2025.

This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.