OJBench: A Competition Level Code Benchmark For Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zhexu, Liu, Yiping, Wang, Yejie, He, Wenyang, Gao, Bofei, Diao, Muxi, Chen, Yanxu, Fu, Kelin, Sung, Flood, Yang, Zhilin, Liu, Tianyu, Xu, Weiran
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908859485913088
author Wang, Zhexu
Liu, Yiping
Wang, Yejie
He, Wenyang
Gao, Bofei
Diao, Muxi
Chen, Yanxu
Fu, Kelin
Sung, Flood
Yang, Zhilin
Liu, Tianyu
Xu, Weiran
author_facet Wang, Zhexu
Liu, Yiping
Wang, Yejie
He, Wenyang
Gao, Bofei
Diao, Muxi
Chen, Yanxu
Fu, Kelin
Sung, Flood
Yang, Zhilin
Liu, Tianyu
Xu, Weiran
contents Recent advancements in large language models (LLMs) have demonstrated significant progress in math and code reasoning capabilities. However, existing code benchmark are limited in their ability to evaluate the full spectrum of these capabilities, particularly at the competitive level. To bridge this gap, we introduce OJBench, a novel and challenging benchmark designed to assess the competitive-level code reasoning abilities of LLMs. OJBench comprises 232 programming competition problems from NOI and ICPC, providing a more rigorous test of models' reasoning skills. We conducted a comprehensive evaluation using OJBench on 37 models, including both closed-source and open-source models, reasoning-oriented and non-reasoning-oriented models. Our results indicate that even state-of-the-art reasoning-oriented models, such as o4-mini and Gemini-2.5-pro-exp, struggle with highly challenging competition-level problems. This highlights the significant challenges that models face in competitive-level code reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16395
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OJBench: A Competition Level Code Benchmark For Large Language Models
Wang, Zhexu
Liu, Yiping
Wang, Yejie
He, Wenyang
Gao, Bofei
Diao, Muxi
Chen, Yanxu
Fu, Kelin
Sung, Flood
Yang, Zhilin
Liu, Tianyu
Xu, Weiran
Computation and Language
Recent advancements in large language models (LLMs) have demonstrated significant progress in math and code reasoning capabilities. However, existing code benchmark are limited in their ability to evaluate the full spectrum of these capabilities, particularly at the competitive level. To bridge this gap, we introduce OJBench, a novel and challenging benchmark designed to assess the competitive-level code reasoning abilities of LLMs. OJBench comprises 232 programming competition problems from NOI and ICPC, providing a more rigorous test of models' reasoning skills. We conducted a comprehensive evaluation using OJBench on 37 models, including both closed-source and open-source models, reasoning-oriented and non-reasoning-oriented models. Our results indicate that even state-of-the-art reasoning-oriented models, such as o4-mini and Gemini-2.5-pro-exp, struggle with highly challenging competition-level problems. This highlights the significant challenges that models face in competitive-level code reasoning.
title OJBench: A Competition Level Code Benchmark For Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.16395