VERINA: Benchmarking Verifiable Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zhe, Yan, Zhengxu, He, Jingxuan, Kasriel, Timothe, Yang, Kaiyu, Song, Dawn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLEVER: A Curated Benchmark for Formally Verified Code Generation
by: Thakur, Amitayush, et al.
Published: (2025)
by: Thakur, Amitayush, et al.
Published: (2025)
Evaluating the Ability of Large Language Models to Generate Verifiable Specifications in VeriFast
by: Fan, Wen, et al.
Published: (2024)
by: Fan, Wen, et al.
Published: (2024)
Intent-aligned Formal Specification Synthesis via Traceable Refinement
by: Ye, Zhe, et al.
Published: (2026)
by: Ye, Zhe, et al.
Published: (2026)
Dafny as Verification-Aware Intermediate Language for Code Generation
by: Li, Yue Chen, et al.
Published: (2025)
by: Li, Yue Chen, et al.
Published: (2025)
VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
by: Dougherty, Quinn, et al.
Published: (2025)
by: Dougherty, Quinn, et al.
Published: (2025)
Agentic Proving for Program Verification
by: Sosso, Alessandro, et al.
Published: (2026)
by: Sosso, Alessandro, et al.
Published: (2026)
Inferring multiple helper Dafny assertions with LLMs
by: Silva, Álvaro, et al.
Published: (2025)
by: Silva, Álvaro, et al.
Published: (2025)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
by: Lu, Jialin, et al.
Published: (2026)
by: Lu, Jialin, et al.
Published: (2026)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Complete the Cycle: Reachability Types with Expressive Cyclic References (Extended Version)
by: Deng, Haotian, et al.
Published: (2025)
by: Deng, Haotian, et al.
Published: (2025)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
Can Language Models Pretend Solvers? Logic Code Simulation with LLMs
by: Chen, Minyu, et al.
Published: (2024)
by: Chen, Minyu, et al.
Published: (2024)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
by: Chen, Simin, et al.
Published: (2024)
by: Chen, Simin, et al.
Published: (2024)
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
by: Dandamudi, Rohit, et al.
Published: (2024)
by: Dandamudi, Rohit, et al.
Published: (2024)
Multi-Threaded Software Model Checking via Parallel Trace Abstraction Refinement
by: Barth, Max, et al.
Published: (2025)
by: Barth, Max, et al.
Published: (2025)
CHCVerif: A Portfolio-Based Solver for Constrained Horn Clauses
by: Dobos-Kovács, Mihály, et al.
Published: (2025)
by: Dobos-Kovács, Mihály, et al.
Published: (2025)
Guidelines for Producing Concise LNT Models, Illustrated with Formal Models of the Algorand Consensus Protocol
by: Garavel, Hubert
Published: (2026)
by: Garavel, Hubert
Published: (2026)
Proceedings Sixth Workshop on Models for Formal Analysis of Real Systems
by: Lang, Frédéric, et al.
Published: (2024)
by: Lang, Frédéric, et al.
Published: (2024)
Customizing Static Analysis using Codesearch
by: Hayoun, Avi, et al.
Published: (2024)
by: Hayoun, Avi, et al.
Published: (2024)
Proceedings of the 12th Workshop on Horn Clauses for Verification and Synthesis
by: De Angelis, Emanuele, et al.
Published: (2025)
by: De Angelis, Emanuele, et al.
Published: (2025)
Contract Usage and Evolution in Android Mobile Applications
by: Ferreira, David R., et al.
Published: (2024)
by: Ferreira, David R., et al.
Published: (2024)
Establishing tool support for a concept DSL
by: Jakobsen, Nikolaj Kühne
Published: (2025)
by: Jakobsen, Nikolaj Kühne
Published: (2025)
An Enumerative Embedding of the Python Type System in ACL2s
by: Xifaras, Samuel, et al.
Published: (2025)
by: Xifaras, Samuel, et al.
Published: (2025)
GPUMC: A Stateless Model Checker for GPU Weak Memory Concurrency
by: Chakraborty, Soham, et al.
Published: (2025)
by: Chakraborty, Soham, et al.
Published: (2025)
GNU Aris: a web application for students
by: Attri, Saksham, et al.
Published: (2025)
by: Attri, Saksham, et al.
Published: (2025)
Proceedings 9th edition of Working Formal Methods Symposium
by: Arusoaie, Andrei, et al.
Published: (2025)
by: Arusoaie, Andrei, et al.
Published: (2025)
Flexible Correct-by-Construction Programming
by: Runge, Tobias, et al.
Published: (2022)
by: Runge, Tobias, et al.
Published: (2022)
Tunable Automation in Automated Program Verification
by: Bai, Alexander Y., et al.
Published: (2025)
by: Bai, Alexander Y., et al.
Published: (2025)
Programming Really Is Simple Mathematics
by: Meyer, Bertrand, et al.
Published: (2025)
by: Meyer, Bertrand, et al.
Published: (2025)
Context-Sensitive Abstract Interpretation of Dynamic Languages
by: Piszcz, Franciszek
Published: (2024)
by: Piszcz, Franciszek
Published: (2024)
Leveraging Large Language Models to Boost Dafny's Developers Productivity
by: Silva, Álvaro, et al.
Published: (2024)
by: Silva, Álvaro, et al.
Published: (2024)
Taming Silent Failures: A Framework for Verifiable AI Reliability
by: Yang, Guan-Yan, et al.
Published: (2025)
by: Yang, Guan-Yan, et al.
Published: (2025)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
Generating Verifiable Chain of Thoughts from Exection-Traces
by: Thakur, Shailja, et al.
Published: (2025)
by: Thakur, Shailja, et al.
Published: (2025)
Dynamic Stability of LLM-Generated Code
by: Rajput, Prateek, et al.
Published: (2025)
by: Rajput, Prateek, et al.
Published: (2025)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
by: Wallraven, Stephan, et al.
Published: (2026)
by: Wallraven, Stephan, et al.
Published: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
Similar Items
-
CLEVER: A Curated Benchmark for Formally Verified Code Generation
by: Thakur, Amitayush, et al.
Published: (2025) -
Evaluating the Ability of Large Language Models to Generate Verifiable Specifications in VeriFast
by: Fan, Wen, et al.
Published: (2024) -
Intent-aligned Formal Specification Synthesis via Traceable Refinement
by: Ye, Zhe, et al.
Published: (2026) -
Dafny as Verification-Aware Intermediate Language for Code Generation
by: Li, Yue Chen, et al.
Published: (2025) -
VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search
by: Brandfonbrener, David, et al.
Published: (2024)