The Path Not Taken: Duality in Reasoning about Program Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Hasanov, Eshgin, Sibat, Md Mahadi Hassan, Karmaker, Santu, Yadavally, Aashish |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026)
by: Weng, Haojun, et al.
Published: (2026)
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
by: Jana, Prithwish, et al.
Published: (2023)
by: Jana, Prithwish, et al.
Published: (2023)
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
by: Li, Caihua, et al.
Published: (2026)
by: Li, Caihua, et al.
Published: (2026)
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)
by: Gunnu, Sravani, et al.
Published: (2025)
LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems
by: Fakih, Mohamad, et al.
Published: (2024)
by: Fakih, Mohamad, et al.
Published: (2024)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
by: Khatchadourian, Raffi, et al.
Published: (2025)
by: Khatchadourian, Raffi, et al.
Published: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
by: Kiashemshaki, Kiana, et al.
Published: (2025)
by: Kiashemshaki, Kiana, et al.
Published: (2025)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
by: Young, Richard J.
Published: (2025)
by: Young, Richard J.
Published: (2025)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
Enhancing LLM Code Generation Capabilities through Test-Driven Development and Code Interpreter
by: Jalil, Sajed, et al.
Published: (2025)
by: Jalil, Sajed, et al.
Published: (2025)
Anka: A Domain-Specific Language for Reliable LLM Code Generation
by: Mazrouei, Saif Khalfan Saif Al
Published: (2025)
by: Mazrouei, Saif Khalfan Saif Al
Published: (2025)
Compressed code: the hidden effects of quantization and distillation on programming tokens
by: Siniaev, Viacheslav, et al.
Published: (2026)
by: Siniaev, Viacheslav, et al.
Published: (2026)
Exploring LLMs for User Story Extraction from Mockups
by: Firmenich, Diego, et al.
Published: (2026)
by: Firmenich, Diego, et al.
Published: (2026)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
by: Abualazm, Raafat, et al.
Published: (2026)
by: Abualazm, Raafat, et al.
Published: (2026)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
by: Mazaheri, Parsa
Published: (2026)
by: Mazaheri, Parsa
Published: (2026)
Engineering A Large Language Model From Scratch
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
by: Mughal, Ali Hassaan, et al.
Published: (2026)
by: Mughal, Ali Hassaan, et al.
Published: (2026)
Revisiting Word Embeddings in the LLM Era
by: Mahajan, Yash, et al.
Published: (2024)
by: Mahajan, Yash, et al.
Published: (2024)
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
by: Yu, Boxi, et al.
Published: (2026)
by: Yu, Boxi, et al.
Published: (2026)
Leveraging Test Driven Development with Large Language Models for Reliable and Verifiable Spreadsheet Code Generation: A Research Framework
by: Thorne, Simon, et al.
Published: (2025)
by: Thorne, Simon, et al.
Published: (2025)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
by: Iscan, Mehmet
Published: (2026)
by: Iscan, Mehmet
Published: (2026)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
by: Petrovic, Nenad, et al.
Published: (2024)
by: Petrovic, Nenad, et al.
Published: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
by: Lebioda, Krzysztof, et al.
Published: (2024)
by: Lebioda, Krzysztof, et al.
Published: (2024)
TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments
by: Sakizli, Furkan
Published: (2026)
by: Sakizli, Furkan
Published: (2026)
Comprehensive Evaluation of Large Language Models on Software Engineering Tasks: A Multi-Task Benchmark
by: Gunawan, Go Frendi, et al.
Published: (2026)
by: Gunawan, Go Frendi, et al.
Published: (2026)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
by: Sharma, Krishna Kumaar
Published: (2025)
by: Sharma, Krishna Kumaar
Published: (2025)
Plan with Code: Comparing approaches for robust NL to DSL generation
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
by: Bassamzadeh, Nastaran, et al.
Published: (2024)
Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
by: Matotek, Kadin, et al.
Published: (2025)
by: Matotek, Kadin, et al.
Published: (2025)
GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
by: Fachada, Nuno, et al.
Published: (2025)
by: Fachada, Nuno, et al.
Published: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair
by: Dinu, Ion George, et al.
Published: (2026)
by: Dinu, Ion George, et al.
Published: (2026)
FinTradeBench: A Financial Reasoning Benchmark for LLMs
by: Agrawal, Yogesh, et al.
Published: (2026)
by: Agrawal, Yogesh, et al.
Published: (2026)
Large Language Models as Software Components: A Taxonomy for LLM-Integrated Applications
by: Weber, Irene
Published: (2024)
by: Weber, Irene
Published: (2024)
LLMORPH: Automated Metamorphic Testing of Large Language Models
by: Cho, Steven, et al.
Published: (2026)
by: Cho, Steven, et al.
Published: (2026)
Similar Items
-
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
by: Weng, Haojun, et al.
Published: (2026) -
Reducing Maintenance Burden in Behaviour-Driven Development: A Paraphrase-Robust Duplicate-Step Detector with a 1.1M-Step Open Benchmark
by: Mughal, Ali Hassaan, et al.
Published: (2026) -
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
by: Jana, Prithwish, et al.
Published: (2023) -
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
by: Li, Caihua, et al.
Published: (2026) -
CIFE: Code Instruction-Following Evaluation
by: Gunnu, Sravani, et al.
Published: (2025)