LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Zihan, Cheng, Zerui, Shen, Zeyu, Zhou, Shang, Liu, Kaiyuan, He, Hansen, Li, Dongruixuan, Wei, Stanley, Hao, Hangyi, Yao, Jianzhu, Sheng, Peiyao, Wang, Zixuan, Chai, Wenhao, Korolova, Aleksandra, Henderson, Peter, Arora, Sanjeev, Viswanath, Pramod, Shang, Jingbo, Xie, Saining |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
by: Cheng, Zerui, et al.
Published: (2026)
by: Cheng, Zerui, et al.
Published: (2026)
VeRA: Verified Reasoning Data Augmentation at Scale
by: Cheng, Zerui, et al.
Published: (2026)
by: Cheng, Zerui, et al.
Published: (2026)
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
by: Yao, Jianzhu, et al.
Published: (2025)
by: Yao, Jianzhu, et al.
Published: (2025)
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation
by: Zhou, Shang, et al.
Published: (2026)
by: Zhou, Shang, et al.
Published: (2026)
Adaptive Curves for Optimally Efficient Market Making
by: Nadkarni, Viraj, et al.
Published: (2024)
by: Nadkarni, Viraj, et al.
Published: (2024)
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
by: Yao, Jianzhu, et al.
Published: (2025)
by: Yao, Jianzhu, et al.
Published: (2025)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
External Evaluation of Discrimination Mitigation Efforts in Meta's Ad Delivery
by: Imana, Basileal, et al.
Published: (2025)
by: Imana, Basileal, et al.
Published: (2025)
Why am I Still Seeing This: Measuring the Effectiveness Of Ad Controls and Explanations in AI-Mediated Ad Targeting Systems
by: Castleman, Jane, et al.
Published: (2024)
by: Castleman, Jane, et al.
Published: (2024)
Adultification Bias in LLMs and Text-to-Image Models
by: Castleman, Jane, et al.
Published: (2025)
by: Castleman, Jane, et al.
Published: (2025)
Proof of Diligence: Cryptoeconomic Security for Rollups
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
by: Irwin, Lucas, et al.
Published: (2025)
by: Irwin, Lucas, et al.
Published: (2025)
BFT-PoLoc: A Byzantine Fortified Trigonometric Proof of Location Protocol using Internet Delays
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
Unconditionally Safe Light Client
by: Moshrefi, Niusha, et al.
Published: (2024)
by: Moshrefi, Niusha, et al.
Published: (2024)
Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
MURMUR: Using cross-user chatter to break collaborative language agents in groups
by: Patlan, Atharv Singh, et al.
Published: (2025)
by: Patlan, Atharv Singh, et al.
Published: (2025)
Measuring Validity in LLM-based Resume Screening
by: Castleman, Jane, et al.
Published: (2026)
by: Castleman, Jane, et al.
Published: (2026)
Auditing for Racial Discrimination in the Delivery of Education Ads
by: Imana, Basileal, et al.
Published: (2024)
by: Imana, Basileal, et al.
Published: (2024)
Auditing for Bias in Ad Delivery Using Inferred Demographic Attributes
by: Imana, Basileal, et al.
Published: (2024)
by: Imana, Basileal, et al.
Published: (2024)
Split the Yield, Share the Risk: Pricing, Hedging and Fixed rates in DeFi
by: Nadkarni, Viraj, et al.
Published: (2025)
by: Nadkarni, Viraj, et al.
Published: (2025)
FrontierCS: Evolving Challenges for Evolving Intelligence
by: Mang, Qiuyang, et al.
Published: (2025)
by: Mang, Qiuyang, et al.
Published: (2025)
Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret with Infinite Variance
by: Zhao, Hangyi
Published: (2026)
by: Zhao, Hangyi
Published: (2026)
Insider Purchase Signals in Microcap Equities: Gradient Boosting Detection of Abnormal Returns
by: Zhao, Hangyi
Published: (2026)
by: Zhao, Hangyi
Published: (2026)
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
by: Shen, Zeyu, et al.
Published: (2025)
by: Shen, Zeyu, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
CFT-Forensics: High-Performance Byzantine Accountability for Crash Fault Tolerant Protocols
by: Tang, Weizhao, et al.
Published: (2023)
by: Tang, Weizhao, et al.
Published: (2023)
Scalable Fingerprinting of Large Language Models
by: Nasery, Anshul, et al.
Published: (2025)
by: Nasery, Anshul, et al.
Published: (2025)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Codifying Character Logic in Role-Playing
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Watermarks for Language Models via Probabilistic Automata
by: Wang, Yangkun, et al.
Published: (2025)
by: Wang, Yangkun, et al.
Published: (2025)
Entangled Relations: Leveraging NLI and Meta-analysis to Enhance Biomedical Relation Extraction
by: Hogan, William, et al.
Published: (2024)
by: Hogan, William, et al.
Published: (2024)
Paremias of the Latvians and the Russians in Latgale: From the Holy Scripture to Modern Existence
by: Jelena Korolova
Published: (2020)
by: Jelena Korolova
Published: (2020)
Thinking Fast and Slow: Data-Driven Adaptive DeFi Borrow-Lending Protocol
by: Bastankhah, Mahsa, et al.
Published: (2024)
by: Bastankhah, Mahsa, et al.
Published: (2024)
ZeroSwap: Data-driven Optimal Market Making in DeFi
by: Nadkarni, Viraj, et al.
Published: (2023)
by: Nadkarni, Viraj, et al.
Published: (2023)
Stability and Multigroup Fairness in Ranking with Uncertain Predictions
by: Devic, Siddartha, et al.
Published: (2024)
by: Devic, Siddartha, et al.
Published: (2024)
On the Use of Proxies in Political Ad Targeting
by: Sapiezynski, Piotr, et al.
Published: (2024)
by: Sapiezynski, Piotr, et al.
Published: (2024)
Virginia Haviland--1976 Regina Medalist--Presentation and Acceptance
by: Field, Carolyn Wicker, et al.
Published: (1976)
by: Field, Carolyn Wicker, et al.
Published: (1976)
Similar Items
-
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025) -
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
by: Cheng, Zerui, et al.
Published: (2026) -
VeRA: Verified Reasoning Data Augmentation at Scale
by: Cheng, Zerui, et al.
Published: (2026) -
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks
by: Yao, Jianzhu, et al.
Published: (2025) -
OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation
by: Zhou, Shang, et al.
Published: (2026)