Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dougherty, Quinn, Mehta, Ronak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
VERINA: Benchmarking Verifiable Code Generation
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis
von: Murphy, William, et al.
Veröffentlicht: (2024)
von: Murphy, William, et al.
Veröffentlicht: (2024)
Taming Silent Failures: A Framework for Verifiable AI Reliability
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2025)
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2025)
Intent-aligned Formal Specification Synthesis via Traceable Refinement
von: Ye, Zhe, et al.
Veröffentlicht: (2026)
von: Ye, Zhe, et al.
Veröffentlicht: (2026)
VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
Can Language Models Pretend Solvers? Logic Code Simulation with LLMs
von: Chen, Minyu, et al.
Veröffentlicht: (2024)
von: Chen, Minyu, et al.
Veröffentlicht: (2024)
Next Steps in LLM-Supported Java Verification
von: Teuber, Samuel, et al.
Veröffentlicht: (2025)
von: Teuber, Samuel, et al.
Veröffentlicht: (2025)
RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2025)
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2025)
RocqSmith: Can Automatic Optimization Forge Better Proof Agents?
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2026)
von: Kozyrev, Andrei, et al.
Veröffentlicht: (2026)
MPBMC: Multi-Property Bounded Model Checking with GNN-guided Clustering
von: Roy, Soumik Guha, et al.
Veröffentlicht: (2026)
von: Roy, Soumik Guha, et al.
Veröffentlicht: (2026)
Agentic Proving for Program Verification
von: Sosso, Alessandro, et al.
Veröffentlicht: (2026)
von: Sosso, Alessandro, et al.
Veröffentlicht: (2026)
A DPLL(T) Framework for Verifying Deep Neural Networks
von: Duong, Hai, et al.
Veröffentlicht: (2023)
von: Duong, Hai, et al.
Veröffentlicht: (2023)
Runtime Monitoring and Enforcement of Conditional Fairness in Generative AIs
von: Cheng, Chih-Hong, et al.
Veröffentlicht: (2024)
von: Cheng, Chih-Hong, et al.
Veröffentlicht: (2024)
StatWhy: Formal Verification Tool for Statistical Hypothesis Testing Programs
von: Kawamoto, Yusuke, et al.
Veröffentlicht: (2024)
von: Kawamoto, Yusuke, et al.
Veröffentlicht: (2024)
LTLGuard: Formalizing LTL Specifications with Compact Language Models and Lightweight Symbolic Reasoning
von: Andresel, Medina, et al.
Veröffentlicht: (2026)
von: Andresel, Medina, et al.
Veröffentlicht: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
Viverra: Text-to-Code with Guarantees
von: Wu, Haoze, et al.
Veröffentlicht: (2026)
von: Wu, Haoze, et al.
Veröffentlicht: (2026)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
von: Lu, Jialin, et al.
Veröffentlicht: (2026)
von: Lu, Jialin, et al.
Veröffentlicht: (2026)
Dafny as Verification-Aware Intermediate Language for Code Generation
von: Li, Yue Chen, et al.
Veröffentlicht: (2025)
von: Li, Yue Chen, et al.
Veröffentlicht: (2025)
Evaluating the Ability of Large Language Models to Generate Verifiable Specifications in VeriFast
von: Fan, Wen, et al.
Veröffentlicht: (2024)
von: Fan, Wen, et al.
Veröffentlicht: (2024)
What are the Right Symmetries for Formal Theorem Proving?
von: Olejniczak, Krzysztof, et al.
Veröffentlicht: (2026)
von: Olejniczak, Krzysztof, et al.
Veröffentlicht: (2026)
Clover: Closed-Loop Verifiable Code Generation
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
Position: Vibe Coding Needs Vibe Reasoning: Improving Vibe Coding with Formal Verification
von: Mitchell, Jacqueline, et al.
Veröffentlicht: (2025)
von: Mitchell, Jacqueline, et al.
Veröffentlicht: (2025)
Context-Augmented Code Generation: How Product Context Improves AI Coding Agent Decision Compliance by 49%
von: Dillon, Drew, et al.
Veröffentlicht: (2026)
von: Dillon, Drew, et al.
Veröffentlicht: (2026)
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
LLMs and Fuzzing in Tandem: A New Approach to Automatically Generating Weakest Preconditions
von: King, Daragh, et al.
Veröffentlicht: (2025)
von: King, Daragh, et al.
Veröffentlicht: (2025)
LeanAgent: Lifelong Learning for Formal Theorem Proving
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2024)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2024)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
MacroSwarm: A Field-based Compositional Framework for Swarm Programming
von: Aguzzi, Gianluca, et al.
Veröffentlicht: (2024)
von: Aguzzi, Gianluca, et al.
Veröffentlicht: (2024)
BAIT: Benchmarking (Embedding) Architectures for Interactive Theorem-Proving
von: Lamont, Sean, et al.
Veröffentlicht: (2024)
von: Lamont, Sean, et al.
Veröffentlicht: (2024)
Correct-by-Construction G-Code Generation: A Neuro-Symbolic Approach via Separation Logic
von: Lee, Yeonseok
Veröffentlicht: (2026)
von: Lee, Yeonseok
Veröffentlicht: (2026)
Watchdogs and Oracles: Runtime Verification Meets Large Language Models for Autonomous Systems
von: Ferrando, Angelo
Veröffentlicht: (2025)
von: Ferrando, Angelo
Veröffentlicht: (2025)
Pseudo-Boolean d-DNNF Compilation for Expressive Feature Modeling Constructs
von: Sundermann, Chico, et al.
Veröffentlicht: (2025)
von: Sundermann, Chico, et al.
Veröffentlicht: (2025)
Lawful and Accountable Personal Data Processing with GDPR-based Access and Usage Control in Distributed Systems
von: van Binsbergen, L. Thomas, et al.
Veröffentlicht: (2025)
von: van Binsbergen, L. Thomas, et al.
Veröffentlicht: (2025)
Accelerating Policy Synthesis in Large-Scale MDPs via Hierarchical Adaptive Refinement
von: Evangelidis, Alexandros, et al.
Veröffentlicht: (2025)
von: Evangelidis, Alexandros, et al.
Veröffentlicht: (2025)
Layered and Staged Monte Carlo Tree Search for SMT Strategy Synthesis
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
Declarative Scenario-based Testing with RoadLogic
von: Bartocci, Ezio, et al.
Veröffentlicht: (2026)
von: Bartocci, Ezio, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025) -
VERINA: Benchmarking Verifiable Code Generation
von: Ye, Zhe, et al.
Veröffentlicht: (2025) -
Combining LLM Code Generation with Formal Specifications and Reactive Program Synthesis
von: Murphy, William, et al.
Veröffentlicht: (2024) -
Taming Silent Failures: A Framework for Verifiable AI Reliability
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2025) -
Intent-aligned Formal Specification Synthesis via Traceable Refinement
von: Ye, Zhe, et al.
Veröffentlicht: (2026)