ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Shiyi, Hu, Yiwen, Min, Yingqian, Chen, Zhipeng, Zhao, Wayne Xin, Wen, Ji-Rong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912415233343488
author Xu, Shiyi
Hu, Yiwen
Min, Yingqian
Chen, Zhipeng
Zhao, Wayne Xin
Wen, Ji-Rong
author_facet Xu, Shiyi
Hu, Yiwen
Min, Yingqian
Chen, Zhipeng
Zhao, Wayne Xin
Wen, Ji-Rong
contents With the significant progress of large reasoning models in complex coding and reasoning tasks, existing benchmarks, like LiveCodeBench and CodeElo, are insufficient to evaluate the coding capabilities of large language models (LLMs) in real competition environments. Moreover, current evaluation metrics such as Pass@K fail to capture the reflective abilities of reasoning models. To address these challenges, we propose \textbf{ICPC-Eval}, a top-level competitive coding benchmark designed to probing the frontiers of LLM reasoning. ICPC-Eval includes 118 carefully curated problems from 11 recent ICPC contests held in various regions of the world, offering three key contributions: 1) A challenging realistic ICPC competition scenario, featuring a problem type and difficulty distribution consistent with actual contests. 2) A robust test case generation method and a corresponding local evaluation toolkit, enabling efficient and accurate local evaluation. 3) An effective test-time scaling evaluation metric, Refine@K, which allows iterative repair of solutions based on execution feedback. The results underscore the significant challenge in evaluating complex reasoning abilities: top-tier reasoning models like DeepSeek-R1 often rely on multi-turn code feedback to fully unlock their in-context reasoning potential when compared to non-reasoning counterparts. Furthermore, despite recent advancements in code generation, these models still lag behind top-performing human teams. We release the benchmark at: https://github.com/RUCAIBox/Slow_Thinking_with_LLMs
format Preprint
id arxiv_https___arxiv_org_abs_2506_04894
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
Xu, Shiyi
Hu, Yiwen
Min, Yingqian
Chen, Zhipeng
Zhao, Wayne Xin
Wen, Ji-Rong
Computation and Language
With the significant progress of large reasoning models in complex coding and reasoning tasks, existing benchmarks, like LiveCodeBench and CodeElo, are insufficient to evaluate the coding capabilities of large language models (LLMs) in real competition environments. Moreover, current evaluation metrics such as Pass@K fail to capture the reflective abilities of reasoning models. To address these challenges, we propose \textbf{ICPC-Eval}, a top-level competitive coding benchmark designed to probing the frontiers of LLM reasoning. ICPC-Eval includes 118 carefully curated problems from 11 recent ICPC contests held in various regions of the world, offering three key contributions: 1) A challenging realistic ICPC competition scenario, featuring a problem type and difficulty distribution consistent with actual contests. 2) A robust test case generation method and a corresponding local evaluation toolkit, enabling efficient and accurate local evaluation. 3) An effective test-time scaling evaluation metric, Refine@K, which allows iterative repair of solutions based on execution feedback. The results underscore the significant challenge in evaluating complex reasoning abilities: top-tier reasoning models like DeepSeek-R1 often rely on multi-turn code feedback to fully unlock their in-context reasoning potential when compared to non-reasoning counterparts. Furthermore, despite recent advancements in code generation, these models still lag behind top-performing human teams. We release the benchmark at: https://github.com/RUCAIBox/Slow_Thinking_with_LLMs
title ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
topic Computation and Language
url https://arxiv.org/abs/2506.04894