Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Xu, Li, Xin, Qu, Xingwei, Fu, Jie, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
by: Zeng, Lingfei, et al.
Published: (2025)
by: Zeng, Lingfei, et al.
Published: (2025)
Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
by: Councilman, Aaron, et al.
Published: (2025)
by: Councilman, Aaron, et al.
Published: (2025)
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
by: Richter, Cedric, et al.
Published: (2025)
by: Richter, Cedric, et al.
Published: (2025)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
by: Liu, Chengwu, et al.
Published: (2025)
by: Liu, Chengwu, et al.
Published: (2025)
Towards Repository-Level Program Verification with Large Language Models
by: Zhong, Si Cheng, et al.
Published: (2025)
by: Zhong, Si Cheng, et al.
Published: (2025)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
by: Endres, Madeline, et al.
Published: (2023)
by: Endres, Madeline, et al.
Published: (2023)
UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models
by: He, Guangxin, et al.
Published: (2025)
by: He, Guangxin, et al.
Published: (2025)
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
by: Dong, Yixin, et al.
Published: (2024)
by: Dong, Yixin, et al.
Published: (2024)
Cobblestone: A Divide-and-Conquer Approach for Automating Formal Verification
by: Kasibatla, Saketh Ram, et al.
Published: (2024)
by: Kasibatla, Saketh Ram, et al.
Published: (2024)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
by: LI, Yizhi, et al.
Published: (2024)
by: LI, Yizhi, et al.
Published: (2024)
Relating Answer Set Programming and Many-sorted Logics for Formal Verification
by: Hansen, Zachary
Published: (2025)
by: Hansen, Zachary
Published: (2025)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
LocalGPT: Benchmarking and Advancing Large Language Models for Local Life Services in Meituan
by: Lan, Xiaochong, et al.
Published: (2025)
by: Lan, Xiaochong, et al.
Published: (2025)
VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
by: Yubeaton, Patrick, et al.
Published: (2025)
by: Yubeaton, Patrick, et al.
Published: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
by: Cummins, Chris, et al.
Published: (2024)
by: Cummins, Chris, et al.
Published: (2024)
A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants
by: Bayazıt, Barış, et al.
Published: (2025)
by: Bayazıt, Barış, et al.
Published: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Observing Micromotives and Macrobehavior of Large Language Models
by: Cheng, Yuyang, et al.
Published: (2024)
by: Cheng, Yuyang, et al.
Published: (2024)
OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification
by: Li, Shangyu, et al.
Published: (2025)
by: Li, Shangyu, et al.
Published: (2025)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
by: Wallraven, Stephan, et al.
Published: (2026)
by: Wallraven, Stephan, et al.
Published: (2026)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
by: Ravi, Nikil, et al.
Published: (2026)
by: Ravi, Nikil, et al.
Published: (2026)
SuperCoder: Assembly Program Superoptimization with Large Language Models
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
Adaptive Draft-Verification for Efficient Large Language Model Decoding
by: Liu, Xukun, et al.
Published: (2024)
by: Liu, Xukun, et al.
Published: (2024)
On the Effectiveness of Large Language Models in Writing Alloy Formulas
by: Hong, Yang, et al.
Published: (2025)
by: Hong, Yang, et al.
Published: (2025)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
by: Zhou, Yefan, et al.
Published: (2025)
by: Zhou, Yefan, et al.
Published: (2025)
Beyond Pass-by-Pass Optimization: Intent-Driven IR Optimization with Large Language Models
by: Qiu, Lei, et al.
Published: (2026)
by: Qiu, Lei, et al.
Published: (2026)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
by: Rabern, Brian, et al.
Published: (2026)
by: Rabern, Brian, et al.
Published: (2026)
LEGO-Compiler: Enhancing Neural Compilation Through Translation Composability
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
Relax: Composable Abstractions for End-to-End Dynamic Machine Learning
by: Lai, Ruihang, et al.
Published: (2023)
by: Lai, Ruihang, et al.
Published: (2023)
Large Language Models Synergize with Automated Machine Learning
by: Xu, Jinglue, et al.
Published: (2024)
by: Xu, Jinglue, et al.
Published: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
by: Qi, Yunjia, et al.
Published: (2025)
by: Qi, Yunjia, et al.
Published: (2025)
JTON: A Token-Efficient JSON Superset with Zen Grid Tabular Encoding for Large Language Models
by: Nandakishore, Gowthamkumar
Published: (2026)
by: Nandakishore, Gowthamkumar
Published: (2026)
Analyzing the Effectiveness of Large Language Models on Text-to-SQL Synthesis
by: Roberson, Richard, et al.
Published: (2024)
by: Roberson, Richard, et al.
Published: (2024)
TaskBench: Benchmarking Large Language Models for Task Automation
by: Shen, Yongliang, et al.
Published: (2023)
by: Shen, Yongliang, et al.
Published: (2023)
Cascading Large Language Models for Salient Event Graph Generation
by: Tan, Xingwei, et al.
Published: (2024)
by: Tan, Xingwei, et al.
Published: (2024)
Similar Items
-
VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable Code
by: Zeng, Lingfei, et al.
Published: (2025) -
Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
by: Councilman, Aaron, et al.
Published: (2025) -
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
by: Richter, Cedric, et al.
Published: (2025) -
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024) -
Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
by: Liu, Chengwu, et al.
Published: (2025)