LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Kaijian, Xiong, Aaron, Zhang, Yunxiang, Zhang, Frederick, Ren, Yueqi, Yang, Jirong, Lee, Ayoung, Bhushan, Shitanshu, Wang, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911330944942080
author Zou, Kaijian
Xiong, Aaron
Zhang, Yunxiang
Zhang, Frederick
Ren, Yueqi
Yang, Jirong
Lee, Ayoung
Bhushan, Shitanshu
Wang, Lu
author_facet Zou, Kaijian
Xiong, Aaron
Zhang, Yunxiang
Zhang, Frederick
Ren, Yueqi
Yang, Jirong
Lee, Ayoung
Bhushan, Shitanshu
Wang, Lu
contents Competitive programming problems increasingly serve as valuable benchmarks to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case coverage, reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a comprehensive benchmark featuring 403 expert-curated Olympiad-level competitive programming problems, each with an average of 60 expert-designed test cases. The problems are sourced directly from 72 official contests of 14 Informatics Olympiads in different regions conducted between 2023 and 2025. LiveOIBench distinguishes itself through four key features: (1) meticulously curated high-quality tasks with detailed subtask rubrics and extensive private test cases; (2) direct integration of elite contestant performance data to enable informative comparison against top-performing humans; (3) planned continuous, contamination-free updates from newly released Olympiad problems; and (4) a self-contained evaluation system facilitating offline and easy-to-reproduce assessments. Benchmarking 34 popular general-purpose and reasoning LLMs, we find that GPT-5 achieves a notable 81.76th percentile, a strong result that nonetheless falls short of top human contestants, who usually place above 90th. In contrast, among open-weight reasoning models, GPT-OSS-120B achieves only a 60th percentile, underscoring significant capability disparities from frontier closed models. Detailed analyses indicate that robust reasoning models prioritize precise problem analysis over excessive exploration, suggesting future models should emphasize structured analysis and minimize unnecessary exploration. All data, code, and leaderboard results are publicly available on our website.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09595
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Zou, Kaijian
Xiong, Aaron
Zhang, Yunxiang
Zhang, Frederick
Ren, Yueqi
Yang, Jirong
Lee, Ayoung
Bhushan, Shitanshu
Wang, Lu
Artificial Intelligence
Computation and Language
Machine Learning
Competitive programming problems increasingly serve as valuable benchmarks to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case coverage, reliance on online platform APIs that limit accessibility. To address these issues, we introduce LiveOIBench, a comprehensive benchmark featuring 403 expert-curated Olympiad-level competitive programming problems, each with an average of 60 expert-designed test cases. The problems are sourced directly from 72 official contests of 14 Informatics Olympiads in different regions conducted between 2023 and 2025. LiveOIBench distinguishes itself through four key features: (1) meticulously curated high-quality tasks with detailed subtask rubrics and extensive private test cases; (2) direct integration of elite contestant performance data to enable informative comparison against top-performing humans; (3) planned continuous, contamination-free updates from newly released Olympiad problems; and (4) a self-contained evaluation system facilitating offline and easy-to-reproduce assessments. Benchmarking 34 popular general-purpose and reasoning LLMs, we find that GPT-5 achieves a notable 81.76th percentile, a strong result that nonetheless falls short of top human contestants, who usually place above 90th. In contrast, among open-weight reasoning models, GPT-OSS-120B achieves only a 60th percentile, underscoring significant capability disparities from frontier closed models. Detailed analyses indicate that robust reasoning models prioritize precise problem analysis over excessive exploration, suggesting future models should emphasize structured analysis and minimize unnecessary exploration. All data, code, and leaderboard results are publicly available on our website.
title LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.09595