AutoCode: LLMs as Problem Setters for Competitive Programming

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Shang, Zheng, Zihan, Liu, Kaiyuan, Shen, Zeyu, Cheng, Zerui, Chen, Zexing, He, Hansen, Yao, Jianzhu, Mao, Huanzhi, Mang, Qiuyang, Fu, Tianfu, Li, Beichen, Li, Dongruixuan, Chai, Wenhao, Liu, Zhuang, Korolova, Aleksandra, Henderson, Peter, Jaques, Natasha, Viswanath, Pramod, Xie, Saining, Shang, Jingbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918160766074880
author Zhou, Shang
Zheng, Zihan
Liu, Kaiyuan
Shen, Zeyu
Cheng, Zerui
Chen, Zexing
He, Hansen
Yao, Jianzhu
Mao, Huanzhi
Mang, Qiuyang
Fu, Tianfu
Li, Beichen
Li, Dongruixuan
Chai, Wenhao
Liu, Zhuang
Korolova, Aleksandra
Henderson, Peter
Jaques, Natasha
Viswanath, Pramod
Xie, Saining
Shang, Jingbo
author_facet Zhou, Shang
Zheng, Zihan
Liu, Kaiyuan
Shen, Zeyu
Cheng, Zerui
Chen, Zexing
He, Hansen
Yao, Jianzhu
Mao, Huanzhi
Mang, Qiuyang
Fu, Tianfu
Li, Beichen
Li, Dongruixuan
Chai, Wenhao
Liu, Zhuang
Korolova, Aleksandra
Henderson, Peter
Jaques, Natasha
Viswanath, Pramod
Xie, Saining
Shang, Jingbo
contents Writing competitive programming problems is exacting. Authors must: set constraints, input distributions, and edge cases that rule out shortcuts; target specific algorithms (e.g., max-flow, dynamic programming, data structures); and calibrate complexity beyond the reach of most competitors. We argue that this makes for an ideal test of general large language model capabilities and study whether they can do this reliably. We introduce AutoCode, which uses multiple rounds of validation to yield competition-grade problem statements and test cases. On held-out problems, AutoCode test suites approach 99% consistency with official judgments, a significant improvement over current state-of-the-art methods like HardTests, which achieve less than 81%. Furthermore, starting with a random seed problem, AutoCode can create novel variants with reference and brute-force solutions. By cross-verifying these generated solutions against test cases, we can further filter out malformed problems. Our system ensures high correctness, as verified by human experts. AutoCode successfully produces novel problems judged by Grandmaster-level (top 0.3%) competitive programmers to be of contest quality.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoCode: LLMs as Problem Setters for Competitive Programming
Zhou, Shang
Zheng, Zihan
Liu, Kaiyuan
Shen, Zeyu
Cheng, Zerui
Chen, Zexing
He, Hansen
Yao, Jianzhu
Mao, Huanzhi
Mang, Qiuyang
Fu, Tianfu
Li, Beichen
Li, Dongruixuan
Chai, Wenhao
Liu, Zhuang
Korolova, Aleksandra
Henderson, Peter
Jaques, Natasha
Viswanath, Pramod
Xie, Saining
Shang, Jingbo
Software Engineering
Artificial Intelligence
Computation and Language
Programming Languages
Writing competitive programming problems is exacting. Authors must: set constraints, input distributions, and edge cases that rule out shortcuts; target specific algorithms (e.g., max-flow, dynamic programming, data structures); and calibrate complexity beyond the reach of most competitors. We argue that this makes for an ideal test of general large language model capabilities and study whether they can do this reliably. We introduce AutoCode, which uses multiple rounds of validation to yield competition-grade problem statements and test cases. On held-out problems, AutoCode test suites approach 99% consistency with official judgments, a significant improvement over current state-of-the-art methods like HardTests, which achieve less than 81%. Furthermore, starting with a random seed problem, AutoCode can create novel variants with reference and brute-force solutions. By cross-verifying these generated solutions against test cases, we can further filter out malformed problems. Our system ensures high correctness, as verified by human experts. AutoCode successfully produces novel problems judged by Grandmaster-level (top 0.3%) competitive programmers to be of contest quality.
title AutoCode: LLMs as Problem Setters for Competitive Programming
topic Software Engineering
Artificial Intelligence
Computation and Language
Programming Languages
url https://arxiv.org/abs/2510.12803