While large language models for code (Code LLMs) excel at generating functionally correct code, existing benchmarks neglect a crucial aspect: adherence to explicit time complexity constraints. We introduce the Complexity-Constraint Code Evaluation (C3E), a novel benchmark evaluating both functional correctness and complexity compliance across feasible and infeasible scenarios. C3E enables precise differentiation between asymptotic complexity classes and tests model robustness against theoretically impossible constraints. Our proposed Complexity Alignment Score (CAS) integrates correctness and complexity adherence into a unified metric, assessed through theoretical analysis rather than costly executions. Experiments reveal a striking gap in state-of-the-art models: GPT-4o achieves 81% correctness but only 31% CAS, demonstrating poor complexity compliance. Notably, most models fail to recognize infeasible constraints except advanced ones such as GPT-4o. These findings underscore the necessity for complexity-aware evaluation, positioning C3E as an essential tool for advancing real-world coding reliability in Code LLMs. The C3E benchmark is available at https://github.com/wahaha12321/C3E.
Publications
- Article type
- Year
- Co-author
Article type
Year
Regular Paper
Issue
Journal of Computer Science and Technology 2026, 41(3): 910-923
Published: 01 May 2026
Total 1
京公网安备11010802044758号