VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Zichen, Pawagi, Mrigank, Liu, Yuxin, Rai, Aaditi, Shao, Lize, Berberian Jr., John, Che, Sicong, Wang, Wenxi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GuardRails: Automated Suggestions for Clarifying Ambiguous Purpose Statements
by: Pawagi, Mrigank, et al.
Published: (2023)
by: Pawagi, Mrigank, et al.
Published: (2023)
AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation
by: Ibrahimzada, Ali Reza, et al.
Published: (2024)
by: Ibrahimzada, Ali Reza, et al.
Published: (2024)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
Probeable Problems for Beginner-level Programming-with-AI Contests
by: Pawagi, Mrigank, et al.
Published: (2024)
by: Pawagi, Mrigank, et al.
Published: (2024)
Property-Driven Evaluation of GNN Expressiveness at Scale: Datasets, Framework, and Study
by: Che, Sicong, et al.
Published: (2026)
by: Che, Sicong, et al.
Published: (2026)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code
by: Awal, Md. Abdul, et al.
Published: (2025)
by: Awal, Md. Abdul, et al.
Published: (2025)
MoEKD: Mixture-of-Experts Knowledge Distillation for Robust and High-Performing Compressed Code Models
by: Awal, Md. Abdul, et al.
Published: (2026)
by: Awal, Md. Abdul, et al.
Published: (2026)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
by: Bai, Yifan, et al.
Published: (2026)
by: Bai, Yifan, et al.
Published: (2026)
VeriFix: Verifying Your Fix Towards An Atomicity Violation
by: Li, Zhuang, et al.
Published: (2025)
by: Li, Zhuang, et al.
Published: (2025)
DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code
by: Parvez, Nabiha, et al.
Published: (2026)
by: Parvez, Nabiha, et al.
Published: (2026)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
by: Awal, Md. Abdul, et al.
Published: (2025)
by: Awal, Md. Abdul, et al.
Published: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
Investigating Adversarial Attacks in Software Analytics via Machine Learning Explainability
by: Awal, MD Abdul, et al.
Published: (2024)
by: Awal, MD Abdul, et al.
Published: (2024)
Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?
by: Awal, Md. Abdul, et al.
Published: (2024)
by: Awal, Md. Abdul, et al.
Published: (2024)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
by: Ebrahimi, Amir M., et al.
Published: (2026)
by: Ebrahimi, Amir M., et al.
Published: (2026)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
Towards Causal Analysis of Empirical Software Engineering Data: The Impact of Programming Languages on Coding Competitions
by: Furia, Carlo A., et al.
Published: (2023)
by: Furia, Carlo A., et al.
Published: (2023)
CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
by: Zhang, Kechi, et al.
Published: (2024)
by: Zhang, Kechi, et al.
Published: (2024)
ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning
by: Tang, Lingxiao, et al.
Published: (2026)
by: Tang, Lingxiao, et al.
Published: (2026)
AetherCode: Evaluating LLMs' Ability to Win In Premier Programming Competitions
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
VeriSBOM: Secure and Verifiable SBOM Sharing Via Zero-Knowledge Proofs
by: Castiglione, Gianpietro, et al.
Published: (2026)
by: Castiglione, Gianpietro, et al.
Published: (2026)
VeriAct: Beyond Verifiability -- Agentic Synthesis of Correct and Complete Formal Specifications
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
Code for All: Educational Applications of the "Vibe Coding" Hackathon in Programming Education across All Skill Levels
by: Chen, Ashley J., et al.
Published: (2026)
by: Chen, Ashley J., et al.
Published: (2026)
ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis
by: Baksys, Mantas, et al.
Published: (2025)
by: Baksys, Mantas, et al.
Published: (2025)
VerifyThisBench: Generating Code, Specifications, and Proofs All at Once
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
An Iterative Test-and-Repair Framework for Competitive Code Generation
by: Tang, Lingxiao, et al.
Published: (2026)
by: Tang, Lingxiao, et al.
Published: (2026)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
by: Miculicich, Lesly, et al.
Published: (2025)
by: Miculicich, Lesly, et al.
Published: (2025)
Benchmarking and Studying the LLM-based Code Review
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
by: Kapadnis, Manav Nitin, et al.
Published: (2025)
TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Towards Verified Code Reasoning by LLMs
by: Sistla, Meghana, et al.
Published: (2025)
by: Sistla, Meghana, et al.
Published: (2025)
Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
by: Lin, Zi, et al.
Published: (2025)
by: Lin, Zi, et al.
Published: (2025)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
by: Wang, Longtian, et al.
Published: (2025)
by: Wang, Longtian, et al.
Published: (2025)
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
by: Xie, Bang, et al.
Published: (2026)
by: Xie, Bang, et al.
Published: (2026)
Human to Document, AI to Code: Comparing GenAI for Notebook Competitions
by: Settewong, Tasha, et al.
Published: (2025)
by: Settewong, Tasha, et al.
Published: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
Similar Items
-
GuardRails: Automated Suggestions for Clarifying Ambiguous Purpose Statements
by: Pawagi, Mrigank, et al.
Published: (2023) -
AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation
by: Ibrahimzada, Ali Reza, et al.
Published: (2024) -
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025) -
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
by: Xie, Zichen, et al.
Published: (2026) -
Probeable Problems for Beginner-level Programming-with-AI Contests
by: Pawagi, Mrigank, et al.
Published: (2024)