Although large language models (LLMs) have been widely deployed across numerous applications, they can generate harmful or illicit content, posing substantial safety risks. Evaluating such risks requires effective evaluation methodologies using high-quality benchmarking datasets. This study introduces LLMSafetyChoice in this regard, a multilingual benchmark for content safety evaluation containing 11911 multiple-choice questions in both Chinese and English, covering four safety domains and eight categories per language (nine in total). We further introduce a systematic metamorphic-testing approach, defining seven metamorphic relations for LLMSafetyChoice, for LLM safety evaluations. Through an extensive empirical study involving 1408 evaluation scenarios (11 LLMs×(8 categories×2 languages)×(1 constructed benchmark + 7 transformations)), we reveal key insights into model behavior under safety-critical conditions and demonstrate that metamorphic testing effectively uncovers subtle safety vulnerabilities. The benchmark and evaluation results are publicly available at https://anonymous.4open.science/r/LLMMetamorphic-08C9/.
Publications
- Article type
- Year
- Co-author
Article type
Year
Regular Paper
Issue
Journal of Computer Science and Technology 2026, 41(3): 924-935
Published: 01 May 2026
Total 1
京公网安备11010802044758号