VerifyThisBench: Generating Code, Specifications, and Proofs All at Once
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Xun, Zhong, Sicheng, Bayazıt, Barış, Veneris, Andreas, Long, Fan, Si, Xujie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing Code Generation with Intermediate Languages
von: Deng, Xun, et al.
Veröffentlicht: (2024)
von: Deng, Xun, et al.
Veröffentlicht: (2024)
RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation
von: Zhong, Sicheng, et al.
Veröffentlicht: (2025)
von: Zhong, Sicheng, et al.
Veröffentlicht: (2025)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
von: Dong, Honghua, et al.
Veröffentlicht: (2025)
Safeguarding DeFi Smart Contracts against Oracle Deviations
von: Deng, Xun, et al.
Veröffentlicht: (2024)
von: Deng, Xun, et al.
Veröffentlicht: (2024)
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
Towards Repository-Level Program Verification with Large Language Models
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
von: Tang, Hao, et al.
Veröffentlicht: (2024)
von: Tang, Hao, et al.
Veröffentlicht: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Enforcing Control Flow Integrity on DeFi Smart Contracts
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
von: Jin, Wen-Long
Veröffentlicht: (2025)
von: Jin, Wen-Long
Veröffentlicht: (2025)
CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
von: Zhang, Kechi, et al.
Veröffentlicht: (2024)
von: Zhang, Kechi, et al.
Veröffentlicht: (2024)
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
von: Saba, Tara, et al.
Veröffentlicht: (2026)
von: Saba, Tara, et al.
Veröffentlicht: (2026)
A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
von: Bayazıt, Barış, et al.
Veröffentlicht: (2025)
von: Bayazıt, Barış, et al.
Veröffentlicht: (2025)
Confidentiality-Preserving Verifiable Business Processes through Zero-Knowledge Proofs
von: Kiesel, Jannis, et al.
Veröffentlicht: (2025)
von: Kiesel, Jannis, et al.
Veröffentlicht: (2025)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
Learning to Solve and Verify: A Self-Play Framework for Code and Test Generation
von: Lin, Zi, et al.
Veröffentlicht: (2025)
von: Lin, Zi, et al.
Veröffentlicht: (2025)
WybeCoder: Verified Imperative Code Generation
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
ExecVerify: White-Box RL with Verifiable Stepwise Rewards for Code Execution Reasoning
von: Tang, Lingxiao, et al.
Veröffentlicht: (2026)
von: Tang, Lingxiao, et al.
Veröffentlicht: (2026)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
Automated Proof Generation for Rust Code via Self-Evolution
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
von: Chen, Tianyu, et al.
Veröffentlicht: (2024)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
From Specifications to Prompts: On the Future of Generative LLMs in Requirements Engineering
von: Vogelsang, Andreas
Veröffentlicht: (2024)
von: Vogelsang, Andreas
Veröffentlicht: (2024)
On the Effectiveness of Large Language Models in Domain-Specific Code Generation
von: Gu, Xiaodong, et al.
Veröffentlicht: (2023)
von: Gu, Xiaodong, et al.
Veröffentlicht: (2023)
From Evaluation to Enhancement: Large Language Models for Zero-Knowledge Proof Code Generation
von: Xue, Zhantong, et al.
Veröffentlicht: (2025)
von: Xue, Zhantong, et al.
Veröffentlicht: (2025)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
von: Li, Haiyang
Veröffentlicht: (2025)
von: Li, Haiyang
Veröffentlicht: (2025)
VeCoGen: Automating Generation of Formally Verified C Code with Large Language Models
von: Sevenhuijsen, Merlijn, et al.
Veröffentlicht: (2024)
von: Sevenhuijsen, Merlijn, et al.
Veröffentlicht: (2024)
Detect Repair Verify for Securing LLM Generated Code: A Multi-Language Empirical Study
von: Cheng, Cheng
Veröffentlicht: (2026)
von: Cheng, Cheng
Veröffentlicht: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
LLMs in Code Vulnerability Analysis: A Proof of Concept
von: Sultana, Shaznin, et al.
Veröffentlicht: (2026)
von: Sultana, Shaznin, et al.
Veröffentlicht: (2026)
Understanding Specification-Driven Code Generation with LLMs: An Empirical Study Design
von: Rosa, Giovanni, et al.
Veröffentlicht: (2026)
von: Rosa, Giovanni, et al.
Veröffentlicht: (2026)
Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
von: Tian, Zhao, et al.
Veröffentlicht: (2023)
von: Tian, Zhao, et al.
Veröffentlicht: (2023)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
Detect--Repair--Verify for LLM-Generated Code: A Multi-Language, Multi-Granularity Empirical Study
von: Cheng, Cheng
Veröffentlicht: (2026)
von: Cheng, Cheng
Veröffentlicht: (2026)
ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis
von: Baksys, Mantas, et al.
Veröffentlicht: (2025)
von: Baksys, Mantas, et al.
Veröffentlicht: (2025)
SAT-DIFF: A Tree Diffing Framework Using SAT Solving
von: Geng, Chuqin, et al.
Veröffentlicht: (2024)
von: Geng, Chuqin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Assessing Code Generation with Intermediate Languages
von: Deng, Xun, et al.
Veröffentlicht: (2024) -
RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation
von: Zhong, Sicheng, et al.
Veröffentlicht: (2025) -
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
von: Dong, Honghua, et al.
Veröffentlicht: (2025) -
Safeguarding DeFi Smart Contracts against Oracle Deviations
von: Deng, Xun, et al.
Veröffentlicht: (2024) -
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
von: Chen, Zhiyang, et al.
Veröffentlicht: (2025)