VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xin, Yutong, Chen, Qiaochu, Durrett, Greg, Dillig, Işil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
von: Wei, Jiayi, et al.
Veröffentlicht: (2023)
von: Wei, Jiayi, et al.
Veröffentlicht: (2023)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025)
DafnyBench: A Benchmark for Formal Software Verification
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
von: Sanchez-Stern, Alex, et al.
Veröffentlicht: (2024)
von: Sanchez-Stern, Alex, et al.
Veröffentlicht: (2024)
SERA: Soft-Verified Efficient Repository Agents
von: Shen, Ethan, et al.
Veröffentlicht: (2026)
von: Shen, Ethan, et al.
Veröffentlicht: (2026)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
von: Zhang, Ke, et al.
Veröffentlicht: (2026)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
Understanding Formal Reasoning Failures in LLMs as Abstract Interpreters
von: Mitchell, Jacqueline L., et al.
Veröffentlicht: (2025)
von: Mitchell, Jacqueline L., et al.
Veröffentlicht: (2025)
Efficient Neural Network Verification via Order Leading Exploration of Branch-and-Bound Trees
von: Zhang, Guanqin, et al.
Veröffentlicht: (2025)
von: Zhang, Guanqin, et al.
Veröffentlicht: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
von: Cao, Jialun, et al.
Veröffentlicht: (2024)
Verification Modulo Tested Library Contracts
von: Uppar, Abhishek, et al.
Veröffentlicht: (2026)
von: Uppar, Abhishek, et al.
Veröffentlicht: (2026)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
A Multi-Perspective Architecture for Semantic Code Search
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2020)
von: Haldar, Rajarshi, et al.
Veröffentlicht: (2020)
Neural Models for Source Code Synthesis and Completion
von: Niyogi, Mitodru
Veröffentlicht: (2024)
von: Niyogi, Mitodru
Veröffentlicht: (2024)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
Translating Large-Scale C Repositories to Idiomatic Rust
von: Dehghan, Saman, et al.
Veröffentlicht: (2025)
von: Dehghan, Saman, et al.
Veröffentlicht: (2025)
FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
von: Chakraborty, Madhurima, et al.
Veröffentlicht: (2025)
von: Chakraborty, Madhurima, et al.
Veröffentlicht: (2025)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
LLMs Lean on Priors, Not Programming Language Semantics
von: Thimmaiah, Aditya, et al.
Veröffentlicht: (2025)
von: Thimmaiah, Aditya, et al.
Veröffentlicht: (2025)
Towards Repository-Level Program Verification with Large Language Models
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
von: Zhong, Si Cheng, et al.
Veröffentlicht: (2025)
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
von: Chen, Boqi, et al.
Veröffentlicht: (2024)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
SaraCoder: Orchestrating Semantic and Structural Cues for Resource-Optimized Repository-Level Code Completion
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaohan, et al.
Veröffentlicht: (2025)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
von: Jain, Naman, et al.
Veröffentlicht: (2024)
von: Jain, Naman, et al.
Veröffentlicht: (2024)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
Linguacodus: A Synergistic Framework for Transformative Code Generation in Machine Learning Pipelines
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
von: Stengel-Eskin, Elias, et al.
Veröffentlicht: (2024)
Guaranteed Guess: A Language Modeling Approach for CISC-to-RISC Transpilation with Testing Guarantees
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback
von: Hajizadeh, Samira, et al.
Veröffentlicht: (2026)
von: Hajizadeh, Samira, et al.
Veröffentlicht: (2026)
Is Programming by Example solved by LLMs?
von: Li, Wen-Ding, et al.
Veröffentlicht: (2024)
von: Li, Wen-Ding, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
von: Wei, Jiayi, et al.
Veröffentlicht: (2023) -
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
von: Khatry, Anirudh, et al.
Veröffentlicht: (2025) -
DafnyBench: A Benchmark for Formal Software Verification
von: Loughridge, Chloe, et al.
Veröffentlicht: (2024) -
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025) -
QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
von: Sanchez-Stern, Alex, et al.
Veröffentlicht: (2024)