Understanding Formal Reasoning Failures in LLMs as Abstract Interpreters
Fuente:
arXiv
Guardado en:
| Autores principales: | Mitchell, Jacqueline L., Kim, Brian Hyeongseok, Zhou, Chenyu, Wang, Chao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FairQuant: Certifying and Quantifying Fairness of Deep Neural Networks
por: Kim, Brian Hyeongseok, et al.
Publicado: (2024)
por: Kim, Brian Hyeongseok, et al.
Publicado: (2024)
Agentic Interpretation: Lattice-Structured Evidence for LLM-Based Program Analysis
por: Mitchell, Jacqueline L., et al.
Publicado: (2026)
por: Mitchell, Jacqueline L., et al.
Publicado: (2026)
FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
por: Chakraborty, Madhurima, et al.
Publicado: (2025)
por: Chakraborty, Madhurima, et al.
Publicado: (2025)
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
por: Zhang, Ke, et al.
Publicado: (2026)
por: Zhang, Ke, et al.
Publicado: (2026)
QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
por: Sanchez-Stern, Alex, et al.
Publicado: (2024)
por: Sanchez-Stern, Alex, et al.
Publicado: (2024)
DINGO: Constrained Inference for Diffusion LLMs
por: Suresh, Tarun, et al.
Publicado: (2025)
por: Suresh, Tarun, et al.
Publicado: (2025)
VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
por: Xin, Yutong, et al.
Publicado: (2026)
por: Xin, Yutong, et al.
Publicado: (2026)
Finding Missed Code Size Optimizations in Compilers using LLMs
por: Italiano, Davide, et al.
Publicado: (2024)
por: Italiano, Davide, et al.
Publicado: (2024)
Teaching LLMs Program Semantics via Symbolic Execution Traces
por: Bayer, Jonas, et al.
Publicado: (2026)
por: Bayer, Jonas, et al.
Publicado: (2026)
Efficient Symbolic Execution of Software under Fault Attacks
por: Fang, Yuzhou, et al.
Publicado: (2025)
por: Fang, Yuzhou, et al.
Publicado: (2025)
An Incremental Algorithm for Algebraic Program Analysis
por: Zhou, Chenyu, et al.
Publicado: (2024)
por: Zhou, Chenyu, et al.
Publicado: (2024)
DafnyBench: A Benchmark for Formal Software Verification
por: Loughridge, Chloe, et al.
Publicado: (2024)
por: Loughridge, Chloe, et al.
Publicado: (2024)
NExT: Teaching Large Language Models to Reason about Code Execution
por: Ni, Ansong, et al.
Publicado: (2024)
por: Ni, Ansong, et al.
Publicado: (2024)
Large Language Models for Multilingual Code Intelligence: A Survey
por: Jiang, Chao, et al.
Publicado: (2026)
por: Jiang, Chao, et al.
Publicado: (2026)
PerfRL: A Small Language Model Framework for Efficient Code Optimization
por: Duan, Shukai, et al.
Publicado: (2023)
por: Duan, Shukai, et al.
Publicado: (2023)
Explainable AI for Embedded Systems Design: A Case Study of Static Redundant NVM Memory Write Prediction
por: Gamatié, Abdoulaye, et al.
Publicado: (2024)
por: Gamatié, Abdoulaye, et al.
Publicado: (2024)
Data Race Detection by Digest-Driven Abstract Interpretation (Extended Version)
por: Schwarz, Michael, et al.
Publicado: (2025)
por: Schwarz, Michael, et al.
Publicado: (2025)
Detecting Buggy Contracts via Smart Testing
por: Wang, Sally Junsong, et al.
Publicado: (2024)
por: Wang, Sally Junsong, et al.
Publicado: (2024)
Is Programming by Example solved by LLMs?
por: Li, Wen-Ding, et al.
Publicado: (2024)
por: Li, Wen-Ding, et al.
Publicado: (2024)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
por: Dai, Hankun, et al.
Publicado: (2025)
por: Dai, Hankun, et al.
Publicado: (2025)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
por: Jiang, Mingchao, et al.
Publicado: (2025)
por: Jiang, Mingchao, et al.
Publicado: (2025)
A benchmark for vericoding: formally verified program synthesis
por: Bursuc, Sergiu, et al.
Publicado: (2025)
por: Bursuc, Sergiu, et al.
Publicado: (2025)
MLCPD: A Unified Multi-Language Code Parsing Dataset with Universal AST Schema
por: Gajjar, Jugal, et al.
Publicado: (2025)
por: Gajjar, Jugal, et al.
Publicado: (2025)
Correctness-Guaranteed Code Generation via Constrained Decoding
por: Li, Lingxiao, et al.
Publicado: (2025)
por: Li, Lingxiao, et al.
Publicado: (2025)
Efficient Neural Network Verification via Order Leading Exploration of Branch-and-Bound Trees
por: Zhang, Guanqin, et al.
Publicado: (2025)
por: Zhang, Guanqin, et al.
Publicado: (2025)
Shedding Light in Task Decomposition in Program Synthesis: The Driving Force of the Synthesizer Model
por: Zenkner, Janis, et al.
Publicado: (2025)
por: Zenkner, Janis, et al.
Publicado: (2025)
Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models
por: Vaidya, Sairam, et al.
Publicado: (2025)
por: Vaidya, Sairam, et al.
Publicado: (2025)
Gradient-Based Program Repair: Fixing Bugs in Continuous Program Spaces
por: Silva, André, et al.
Publicado: (2025)
por: Silva, André, et al.
Publicado: (2025)
GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2
por: Kashmira, Savini, et al.
Publicado: (2025)
por: Kashmira, Savini, et al.
Publicado: (2025)
Automatically Testing Functional Properties of Code Translation Models
por: Eniser, Hasan Ferit, et al.
Publicado: (2023)
por: Eniser, Hasan Ferit, et al.
Publicado: (2023)
SEVerA: Verified Synthesis of Self-Evolving Agents
por: Banerjee, Debangshu, et al.
Publicado: (2026)
por: Banerjee, Debangshu, et al.
Publicado: (2026)
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
por: Chen, Boqi, et al.
Publicado: (2024)
por: Chen, Boqi, et al.
Publicado: (2024)
Coeditor: Leveraging Contextual Changes for Multi-round Code Auto-editing
por: Wei, Jiayi, et al.
Publicado: (2023)
por: Wei, Jiayi, et al.
Publicado: (2023)
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
por: Ugare, Shubham, et al.
Publicado: (2024)
por: Ugare, Shubham, et al.
Publicado: (2024)
JaxDecompiler: Redefining Gradient-Informed Software Design
por: Pochelu, Pierrick
Publicado: (2024)
por: Pochelu, Pierrick
Publicado: (2024)
Language Models for Code Completion: A Practical Evaluation
por: Izadi, Maliheh, et al.
Publicado: (2024)
por: Izadi, Maliheh, et al.
Publicado: (2024)
Constrained Decoding for Fill-in-the-Middle Code Language Models via Efficient Left and Right Quotienting of Context-Sensitive Grammars
por: Melcer, Daniel, et al.
Publicado: (2024)
por: Melcer, Daniel, et al.
Publicado: (2024)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
por: Li, Jingyao, et al.
Publicado: (2023)
por: Li, Jingyao, et al.
Publicado: (2023)
Grounding Data Science Code Generation with Input-Output Specifications
por: Wen, Yeming, et al.
Publicado: (2024)
por: Wen, Yeming, et al.
Publicado: (2024)
WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language Models
por: Yang, Chenyuan, et al.
Publicado: (2023)
por: Yang, Chenyuan, et al.
Publicado: (2023)
Ejemplares similares
-
FairQuant: Certifying and Quantifying Fairness of Deep Neural Networks
por: Kim, Brian Hyeongseok, et al.
Publicado: (2024) -
Agentic Interpretation: Lattice-Structured Evidence for LLM-Based Program Analysis
por: Mitchell, Jacqueline L., et al.
Publicado: (2026) -
FormalSpecCpp: A Dataset of C++ Formal Specifications created using LLMs
por: Chakraborty, Madhurima, et al.
Publicado: (2025) -
Understanding Tool-Augmented Agents for Lean Formalization: A Factorial Analysis
por: Zhang, Ke, et al.
Publicado: (2026) -
QEDCartographer: Automating Formal Verification Using Reward-Free Reinforcement Learning
por: Sanchez-Stern, Alex, et al.
Publicado: (2024)