LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Kaijian, Xiong, Aaron, Zhang, Yunxiang, Zhang, Frederick, Ren, Yueqi, Yang, Jirong, Lee, Ayoung, Bhushan, Shitanshu, Wang, Lu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
by: Zhu, Yaoming, et al.
Published: (2025)
by: Zhu, Yaoming, et al.
Published: (2025)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
On Many-Shot In-Context Learning for Long-Context Evaluation
by: Zou, Kaijian, et al.
Published: (2024)
by: Zou, Kaijian, et al.
Published: (2024)
Can Language Models Solve Olympiad Programming?
by: Shi, Quan, et al.
Published: (2024)
by: Shi, Quan, et al.
Published: (2024)
Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
by: Majmudar, Neh, et al.
Published: (2025)
by: Majmudar, Neh, et al.
Published: (2025)
Can AI Outperform Human Experts in Creating Social Media Creatives?
by: Park, Eunkyung, et al.
Published: (2024)
by: Park, Eunkyung, et al.
Published: (2024)
Mastering Olympiad-Level Physics with Artificial Intelligence
by: Jian, Dong-Shan, et al.
Published: (2025)
by: Jian, Dong-Shan, et al.
Published: (2025)
Linguistics Olympiad
by: Neacșu, Vlad A.
Published: (2024)
by: Neacșu, Vlad A.
Published: (2024)
Proving Olympiad Algebraic Inequalities without Human Demonstrations
by: Wei, Chenrui, et al.
Published: (2024)
by: Wei, Chenrui, et al.
Published: (2024)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
by: Gao, Bofei, et al.
Published: (2024)
by: Gao, Bofei, et al.
Published: (2024)
A sufficient condition for the height function to be constant in $ I_g\times_ρ\mathbb{P}^n $
by: Cao, Kaijian
Published: (2024)
by: Cao, Kaijian
Published: (2024)
Clinical Median Images and Deep Learning: Advancing Automated Detection of Ultrasound Transducer Uniformity Artifacts
by: Yang Kaijian
Published: (2026)
by: Yang Kaijian
Published: (2026)
Transcendence: Generative Models Can Outperform The Experts That Train Them
by: Zhang, Edwin, et al.
Published: (2024)
by: Zhang, Edwin, et al.
Published: (2024)
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
by: He, Chaoqun, et al.
Published: (2024)
by: He, Chaoqun, et al.
Published: (2024)
When Less Is More: Binary Feedback Can Outperform Ordinal Comparisons in Ranking Recovery
by: Xu, Shirong, et al.
Published: (2025)
by: Xu, Shirong, et al.
Published: (2025)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
by: Zheng, Zihan, et al.
Published: (2025)
by: Zheng, Zihan, et al.
Published: (2025)
InsertGNN: Can Graph Neural Networks Outperform Humans in TOEFL Sentence Insertion Problem?
by: Wu, Fang, et al.
Published: (2021)
by: Wu, Fang, et al.
Published: (2021)
Can ChatGPT Outperform Humans in Faking a Personality Assessment While Avoiding Detection?
by: Chet Robie, et al.
Published: (2025)
by: Chet Robie, et al.
Published: (2025)
Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
by: Van Veen, Dave, et al.
Published: (2023)
by: Van Veen, Dave, et al.
Published: (2023)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
by: Cherian, Anoop, et al.
Published: (2024)
by: Cherian, Anoop, et al.
Published: (2024)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
by: Cui, Yiming, et al.
Published: (2025)
by: Cui, Yiming, et al.
Published: (2025)
Can a Single Tree Outperform an Entire Forest?
by: Mao, Qiangqiang, et al.
Published: (2024)
by: Mao, Qiangqiang, et al.
Published: (2024)
Can Competition Outperform Collaboration? The Role of Misbehaving Agents
by: Ballotta, Luca, et al.
Published: (2022)
by: Ballotta, Luca, et al.
Published: (2022)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning
by: Liu, Bing, et al.
Published: (2025)
by: Liu, Bing, et al.
Published: (2025)
CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models
by: Chen, Zhuofan, et al.
Published: (2025)
by: Chen, Zhuofan, et al.
Published: (2025)
Audio Outperforms Text for Visual Decoding
by: Zhang, Zhengdi, et al.
Published: (2026)
by: Zhang, Zhengdi, et al.
Published: (2026)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
by: Jeong, Seunghoon, et al.
Published: (2026)
by: Jeong, Seunghoon, et al.
Published: (2026)
Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning
by: Li, Zenan, et al.
Published: (2025)
by: Li, Zenan, et al.
Published: (2025)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
by: Agrawal, Lakshya A, et al.
Published: (2025)
by: Agrawal, Lakshya A, et al.
Published: (2025)
Language Models Coupled with Metacognition Can Outperform Reasoning Models
by: Khandelwal, Vedant, et al.
Published: (2025)
by: Khandelwal, Vedant, et al.
Published: (2025)
Can Decentralized Control Outperform Centralized? The Role of Communication Latency
by: Ballotta, Luca, et al.
Published: (2021)
by: Ballotta, Luca, et al.
Published: (2021)
Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure
by: Powdthavee, Nattavudh
Published: (2026)
by: Powdthavee, Nattavudh
Published: (2026)
Multi-stage Large Language Model Pipelines Can Outperform GPT-4o in Relevance Assessment
by: Schnabel, Julian A., et al.
Published: (2025)
by: Schnabel, Julian A., et al.
Published: (2025)
DOoM: Difficult Olympiads of Math
by: Kuleshov, Ilya, et al.
Published: (2025)
by: Kuleshov, Ilya, et al.
Published: (2025)
Extreme Points and Large Contests
by: Bolgè, Giovanni Valvassori
Published: (2026)
by: Bolgè, Giovanni Valvassori
Published: (2026)
Similar Items
-
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
by: Zhu, Yaoming, et al.
Published: (2025) -
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025) -
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025) -
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025) -
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025)