AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

CodeRankEval: Benchmarking and Analyzing LLM Performance for Code Ranking

School of Software and Microelectronics, Peking University, Beijing 100871, China
Shenzhen International Graduate School, Tsinghua University, Shenzhen 518000, China
Show Author Information

Abstract

Large language models (LLMs) are increasingly applied across diverse software engineering tasks. Consequently, their ability to effectively rank code quality is crucial for applications like selecting optimal solutions and aiding code review. However, evaluating this essential code ranking capability is hampered by a lack of benchmarks covering diverse paradigms and robustness testing. To address this, we introduce CodeRankEval, a benchmark suite for multi-paradigm evaluation, and CodeRankEval-Perturbed for robustness testing against common code flaws. Our empirical study reveals key insights: pairwise ranking yields the highest accuracy but is costly; listwise is the cheapest and shows comparable performance with pairwise; pointwise generally exhibits lower performance with intermediate cost. Besides, ranking ability correlates positively with generation ability, models show reasonable robustness to perturbations but may exhibit positional bias. Overall, this work provides valuable resources and insights for understanding and improving LLM-based code ranking evaluation.

Electronic Supplementary Material

Download File(s)
JCST-2505-15514-Highlights.pdf (350.9 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 1220-1233

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Chen L-G, Xiao Z, Xu Y-J, et al. CodeRankEval: Benchmarking and Analyzing LLM Performance for Code Ranking. Journal of Computer Science and Technology, 2025, 40(5): 1220-1233. https://doi.org/10.1007/s11390-025-5514-9

1767

Views

1

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 01 May 2025
Accepted: 21 August 2025
Published: 10 September 2025
© Institute of Computing Technology, Chinese Academy of Sciences 2025