AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Complexity-Constraint Code Evaluation: A Benchmark for Time Complexity Compliance in LLM-Generated Code

School of Software and Microelectronics, Peking University, Beijing 100871, China
Institute of Advanced Computing Technology, School of Computer Science and Engineering, Beihang University Beijing 100191, China
Shenzhen International Graduate School, Tsinghua University, Shenzhen 518000, China
Show Author Information

Abstract

While large language models for code (Code LLMs) excel at generating functionally correct code, existing benchmarks neglect a crucial aspect: adherence to explicit time complexity constraints. We introduce the Complexity-Constraint Code Evaluation (C3E), a novel benchmark evaluating both functional correctness and complexity compliance across feasible and infeasible scenarios. C3E enables precise differentiation between asymptotic complexity classes and tests model robustness against theoretically impossible constraints. Our proposed Complexity Alignment Score (CAS) integrates correctness and complexity adherence into a unified metric, assessed through theoretical analysis rather than costly executions. Experiments reveal a striking gap in state-of-the-art models: GPT-4o achieves 81% correctness but only 31% CAS, demonstrating poor complexity compliance. Notably, most models fail to recognize infeasible constraints except advanced ones such as GPT-4o. These findings underscore the necessity for complexity-aware evaluation, positioning C3E as an essential tool for advancing real-world coding reliability in Code LLMs. The C3E benchmark is available at https://github.com/wahaha12321/C3E.

Electronic Supplementary Material

Download File(s)
JCST-2505-15518-Highlights.pdf (364.4 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 910-923

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Chen L-G, Wang X, Chen J-Y, et al. Complexity-Constraint Code Evaluation: A Benchmark for Time Complexity Compliance in LLM-Generated Code. Journal of Computer Science and Technology, 2026, 41(3): 910-923. https://doi.org/10.1007/s11390-025-5518-5

11

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 01 May 2025
Accepted: 30 December 2025
Published: 01 May 2026
© Institute of Computing Technology, Chinese Academy of Sciences 2026