From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bayat, Farima Fatahi, Pezeshkpour, Pouya, Hruschka, Estevam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024)
by: Pezeshkpour, Pouya, et al.
Published: (2024)
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Leveraging Large Language Models to Boost Dafny's Developers Productivity
by: Silva, Álvaro, et al.
Published: (2024)
by: Silva, Álvaro, et al.
Published: (2024)
PROMISE: Proof Automation as Structural Imitation of Human Reasoning
by: Ahn, Youngjoo, et al.
Published: (2026)
by: Ahn, Youngjoo, et al.
Published: (2026)
Combining Logic with Large Language Models for Automatic Debugging and Repair of ASP Programs
by: Brancas, Ricardo, et al.
Published: (2024)
by: Brancas, Ricardo, et al.
Published: (2024)
From Task Solving to Robust Real-World Adaptation in LLM Agents
by: Pezeshkpour, Pouya, et al.
Published: (2026)
by: Pezeshkpour, Pouya, et al.
Published: (2026)
Programming Really Is Simple Mathematics
by: Meyer, Bertrand, et al.
Published: (2025)
by: Meyer, Bertrand, et al.
Published: (2025)
Flexible Correct-by-Construction Programming
by: Runge, Tobias, et al.
Published: (2022)
by: Runge, Tobias, et al.
Published: (2022)
Tunable Automation in Automated Program Verification
by: Bai, Alexander Y., et al.
Published: (2025)
by: Bai, Alexander Y., et al.
Published: (2025)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Insight-RAG: Enhancing LLMs with Insight-Driven Augmentation
by: Pezeshkpour, Pouya, et al.
Published: (2025)
by: Pezeshkpour, Pouya, et al.
Published: (2025)
Context-Sensitive Abstract Interpretation of Dynamic Languages
by: Piszcz, Franciszek
Published: (2024)
by: Piszcz, Franciszek
Published: (2024)
Evaluating the Ability of Large Language Models to Generate Verifiable Specifications in VeriFast
by: Fan, Wen, et al.
Published: (2024)
by: Fan, Wen, et al.
Published: (2024)
Guidelines for Producing Concise LNT Models, Illustrated with Formal Models of the Algorand Consensus Protocol
by: Garavel, Hubert
Published: (2026)
by: Garavel, Hubert
Published: (2026)
Formal Analysis of the Sigmoid Function and Formal Proof of the Universal Approximation Theorem
by: Bryant, Dustin, et al.
Published: (2025)
by: Bryant, Dustin, et al.
Published: (2025)
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
Proceedings Sixth Workshop on Models for Formal Analysis of Real Systems
by: Lang, Frédéric, et al.
Published: (2024)
by: Lang, Frédéric, et al.
Published: (2024)
GPUMC: A Stateless Model Checker for GPU Weak Memory Concurrency
by: Chakraborty, Soham, et al.
Published: (2025)
by: Chakraborty, Soham, et al.
Published: (2025)
Multi-Threaded Software Model Checking via Parallel Trace Abstraction Refinement
by: Barth, Max, et al.
Published: (2025)
by: Barth, Max, et al.
Published: (2025)
Quantifying Software Correctness by Combining Architecture Modeling and Formal Program Analysis
by: Lanzinger, Florian, et al.
Published: (2024)
by: Lanzinger, Florian, et al.
Published: (2024)
A Tale of 1001 LoC: Potential Runtime Error-Guided Specification Synthesis for Verifying Large-Scale Programs
by: Wang, Zhongyi, et al.
Published: (2025)
by: Wang, Zhongyi, et al.
Published: (2025)
Towards Automatic Transformations of Coq Proof Scripts
by: Magaud, Nicolas
Published: (2024)
by: Magaud, Nicolas
Published: (2024)
Agentic Proving for Program Verification
by: Sosso, Alessandro, et al.
Published: (2026)
by: Sosso, Alessandro, et al.
Published: (2026)
AutoDeduct: A Tool for Automated Deductive Verification of C Code
by: Amilon, Jesper, et al.
Published: (2025)
by: Amilon, Jesper, et al.
Published: (2025)
Applying Formal Methods Tools to an Electronic Warfare Codebase (Experience report)
by: Li, Letitia W., et al.
Published: (2026)
by: Li, Letitia W., et al.
Published: (2026)
VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
by: Lu, Jialin, et al.
Published: (2026)
by: Lu, Jialin, et al.
Published: (2026)
Traq: Estimating the Quantum Cost of Classical Programs
by: Peduri, Anurudh, et al.
Published: (2025)
by: Peduri, Anurudh, et al.
Published: (2025)
CHCVerif: A Portfolio-Based Solver for Constrained Horn Clauses
by: Dobos-Kovács, Mihály, et al.
Published: (2025)
by: Dobos-Kovács, Mihály, et al.
Published: (2025)
Proceedings of the 12th Workshop on Horn Clauses for Verification and Synthesis
by: De Angelis, Emanuele, et al.
Published: (2025)
by: De Angelis, Emanuele, et al.
Published: (2025)
Complete the Cycle: Reachability Types with Expressive Cyclic References (Extended Version)
by: Deng, Haotian, et al.
Published: (2025)
by: Deng, Haotian, et al.
Published: (2025)
Establishing tool support for a concept DSL
by: Jakobsen, Nikolaj Kühne
Published: (2025)
by: Jakobsen, Nikolaj Kühne
Published: (2025)
An Enumerative Embedding of the Python Type System in ACL2s
by: Xifaras, Samuel, et al.
Published: (2025)
by: Xifaras, Samuel, et al.
Published: (2025)
GNU Aris: a web application for students
by: Attri, Saksham, et al.
Published: (2025)
by: Attri, Saksham, et al.
Published: (2025)
Proceedings 9th edition of Working Formal Methods Symposium
by: Arusoaie, Andrei, et al.
Published: (2025)
by: Arusoaie, Andrei, et al.
Published: (2025)
Customizing Static Analysis using Codesearch
by: Hayoun, Avi, et al.
Published: (2024)
by: Hayoun, Avi, et al.
Published: (2024)
Contract Usage and Evolution in Android Mobile Applications
by: Ferreira, David R., et al.
Published: (2024)
by: Ferreira, David R., et al.
Published: (2024)
StatWhy: Formal Verification Tool for Statistical Hypothesis Testing Programs
by: Kawamoto, Yusuke, et al.
Published: (2024)
by: Kawamoto, Yusuke, et al.
Published: (2024)
Separation Logic for Verifying Physical Collisions of CNC Programs
by: Lee, Yeonseok
Published: (2026)
by: Lee, Yeonseok
Published: (2026)
Similar Items
-
Multi-Conditional Ranking with Large Language Models
by: Pezeshkpour, Pouya, et al.
Published: (2024) -
AutoPyVerifier: Learning Compact Executable Verifiers for Large Language Model Outputs
by: Pezeshkpour, Pouya, et al.
Published: (2026) -
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
by: Wan, Yuxuan, et al.
Published: (2024) -
Leveraging Large Language Models to Boost Dafny's Developers Productivity
by: Silva, Álvaro, et al.
Published: (2024) -
PROMISE: Proof Automation as Structural Imitation of Human Reasoning
by: Ahn, Youngjoo, et al.
Published: (2026)