AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Safety Evaluation for Large Language Models with Metamorphic Testing: An Empirical Study

School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications Beijing 100876, China
Inspur Cloud Information Technology Co., Ltd., Beijing 100193, China
NSFOCUS Technologies Group Co., Ltd., Beijing 100089, China
Show Author Information

Abstract

Although large language models (LLMs) have been widely deployed across numerous applications, they can generate harmful or illicit content, posing substantial safety risks. Evaluating such risks requires effective evaluation methodologies using high-quality benchmarking datasets. This study introduces LLMSafetyChoice in this regard, a multilingual benchmark for content safety evaluation containing 11911 multiple-choice questions in both Chinese and English, covering four safety domains and eight categories per language (nine in total). We further introduce a systematic metamorphic-testing approach, defining seven metamorphic relations for LLMSafetyChoice, for LLM safety evaluations. Through an extensive empirical study involving 1408 evaluation scenarios (11 LLMs×(8 categories×2 languages)×(1 constructed benchmark + 7 transformations)), we reveal key insights into model behavior under safety-critical conditions and demonstrate that metamorphic testing effectively uncovers subtle safety vulnerabilities. The benchmark and evaluation results are publicly available at https://anonymous.4open.science/r/LLMMetamorphic-08C9/.

Electronic Supplementary Material

Download File(s)
JCST-2507-15705-Highlights.pdf (179.9 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 924-935

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Xing Y, Huang J-Q, Zhao H-J, et al. Safety Evaluation for Large Language Models with Metamorphic Testing: An Empirical Study. Journal of Computer Science and Technology, 2026, 41(3): 924-935. https://doi.org/10.1007/s11390-026-5705-z

10

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 04 July 2025
Accepted: 14 May 2026
Published: 01 May 2026
© Institute of Computing Technology, Chinese Academy of Sciences 2026