Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Le-Cong, Thanh, Le, Bach, Murray, Toby |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Memory-Efficient Large Language Models for Program Repair with Semantic-Guided Patch Generation
par: Le-Cong, Thanh, et autres
Publié: (2024)
par: Le-Cong, Thanh, et autres
Publié: (2024)
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
par: Lang, Nguyet-Anh H., et autres
Publié: (2026)
par: Lang, Nguyet-Anh H., et autres
Publié: (2026)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
par: Le-Cong, Thanh, et autres
Publié: (2024)
par: Le-Cong, Thanh, et autres
Publié: (2024)
Can LLMs Enable Verification in Mainstream Programming?
par: Shefer, Aleksandr, et autres
Publié: (2025)
par: Shefer, Aleksandr, et autres
Publié: (2025)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
par: Ho, Anh, et autres
Publié: (2025)
par: Ho, Anh, et autres
Publié: (2025)
LLMs Lean on Priors, Not Programming Language Semantics
par: Thimmaiah, Aditya, et autres
Publié: (2025)
par: Thimmaiah, Aditya, et autres
Publié: (2025)
FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
par: Chakraborty, Madhurima, et autres
Publié: (2025)
par: Chakraborty, Madhurima, et autres
Publié: (2025)
Doc2Spec: Synthesizing Formal Programming Specifications from Natural Language via Grammar Induction
par: Xia, Shihao, et autres
Publié: (2026)
par: Xia, Shihao, et autres
Publié: (2026)
Structured Program Synthesis using LLMs: Results and Insights from the IPARC Challenge
par: Surana, Shraddha, et autres
Publié: (2025)
par: Surana, Shraddha, et autres
Publié: (2025)
AutoCode: LLMs as Problem Setters for Competitive Programming
par: Zhou, Shang, et autres
Publié: (2025)
par: Zhou, Shang, et autres
Publié: (2025)
BODHI: Precise OS Kernel Specification Inference
par: Chang, Zhiming, et autres
Publié: (2026)
par: Chang, Zhiming, et autres
Publié: (2026)
Assessing Code Understanding in LLMs
par: Laneve, Cosimo, et autres
Publié: (2025)
par: Laneve, Cosimo, et autres
Publié: (2025)
Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
par: Endres, Madeline, et autres
Publié: (2023)
par: Endres, Madeline, et autres
Publié: (2023)
Is Programming by Example solved by LLMs?
par: Li, Wen-Ding, et autres
Publié: (2024)
par: Li, Wen-Ding, et autres
Publié: (2024)
A Tool for Automated Reasoning About Traces Based on Configurable Formal Semantics
par: Erata, Ferhat, et autres
Publié: (2024)
par: Erata, Ferhat, et autres
Publié: (2024)
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
par: Richter, Cedric, et autres
Publié: (2025)
par: Richter, Cedric, et autres
Publié: (2025)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
par: Yang, Hua, et autres
Publié: (2025)
par: Yang, Hua, et autres
Publié: (2025)
The New Compiler Stack: A Survey on the Synergy of LLMs and Compilers
par: Zhang, Shuoming, et autres
Publié: (2026)
par: Zhang, Shuoming, et autres
Publié: (2026)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
par: Wang, Ming, et autres
Publié: (2024)
par: Wang, Ming, et autres
Publié: (2024)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
par: Zhang, Yinger, et autres
Publié: (2023)
par: Zhang, Yinger, et autres
Publié: (2023)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
par: Fang, Sen, et autres
Publié: (2025)
par: Fang, Sen, et autres
Publié: (2025)
Raw Pointer Rewriting with LLMs for Translating C to Safer Rust
par: Gao, Yifei, et autres
Publié: (2025)
par: Gao, Yifei, et autres
Publié: (2025)
OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification
par: Li, Shangyu, et autres
Publié: (2025)
par: Li, Shangyu, et autres
Publié: (2025)
Leveraging LLMs to support co-evolution between definitions and instances of textual DSLs
par: Zhang, Weixing, et autres
Publié: (2025)
par: Zhang, Weixing, et autres
Publié: (2025)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
par: Kim, Su-Hyeon, et autres
Publié: (2025)
par: Kim, Su-Hyeon, et autres
Publié: (2025)
Evaluating Program Semantics Reasoning with Type Inference in System F
par: He, Yifeng, et autres
Publié: (2025)
par: He, Yifeng, et autres
Publié: (2025)
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
par: Dong, Honghua, et autres
Publié: (2025)
par: Dong, Honghua, et autres
Publié: (2025)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
par: Tang, Hao, et autres
Publié: (2024)
par: Tang, Hao, et autres
Publié: (2024)
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
par: Le-Cong, Thanh, et autres
Publié: (2026)
par: Le-Cong, Thanh, et autres
Publié: (2026)
Agentic Specification Generator for Move Programs
par: Fu, Yu-Fu, et autres
Publié: (2025)
par: Fu, Yu-Fu, et autres
Publié: (2025)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
par: Agarwal, Anmol, et autres
Publié: (2026)
par: Agarwal, Anmol, et autres
Publié: (2026)
Evaluating the Performance of Large Language Models in Competitive Programming: A Multi-Year, Multi-Grade Analysis
par: Dumitran, Adrian Marius, et autres
Publié: (2024)
par: Dumitran, Adrian Marius, et autres
Publié: (2024)
Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace
par: Yu, Simon, et autres
Publié: (2026)
par: Yu, Simon, et autres
Publié: (2026)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
par: Chen, Simin, et autres
Publié: (2024)
par: Chen, Simin, et autres
Publié: (2024)
Program Skeletons for Automated Program Translation
par: Wang, Bo, et autres
Publié: (2025)
par: Wang, Bo, et autres
Publié: (2025)
Specification-Guided Repair of Arithmetic Errors in Dafny Programs using LLMs
par: Wu, Valentina, et autres
Publié: (2025)
par: Wu, Valentina, et autres
Publié: (2025)
Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
par: Khan, M Zafir Sadik, et autres
Publié: (2026)
par: Khan, M Zafir Sadik, et autres
Publié: (2026)
EnvTrace: Simulation-Based Semantic Evaluation of LLM Code via Execution Trace Alignment -- Demonstrated at Synchrotron Beamlines
par: van der Vleuten, Noah, et autres
Publié: (2025)
par: van der Vleuten, Noah, et autres
Publié: (2025)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
par: Cao, Yuhan, et autres
Publié: (2025)
par: Cao, Yuhan, et autres
Publié: (2025)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
par: Wei, Anjiang, et autres
Publié: (2025)
par: Wei, Anjiang, et autres
Publié: (2025)
Documents similaires
-
Memory-Efficient Large Language Models for Program Repair with Semantic-Guided Patch Generation
par: Le-Cong, Thanh, et autres
Publié: (2024) -
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
par: Lang, Nguyet-Anh H., et autres
Publié: (2026) -
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
par: Le-Cong, Thanh, et autres
Publié: (2024) -
Can LLMs Enable Verification in Mainstream Programming?
par: Shefer, Aleksandr, et autres
Publié: (2025) -
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
par: Ho, Anh, et autres
Publié: (2025)