Disproving Program Equivalence with LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Allamanis, Miltiadis, Yin, Pengcheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
von: Ni, Ansong, et al.
Veröffentlicht: (2024)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
von: Hooda, Ashish, et al.
Veröffentlicht: (2024)
CigaR: Cost-efficient Program Repair with LLMs
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
LLMs as Compiler for Arabic Programming Language
von: Sibaee, Serry, et al.
Veröffentlicht: (2024)
von: Sibaee, Serry, et al.
Veröffentlicht: (2024)
Teaching LLMs Program Semantics via Symbolic Execution Traces
von: Bayer, Jonas, et al.
Veröffentlicht: (2026)
von: Bayer, Jonas, et al.
Veröffentlicht: (2026)
Teaching Code Refactoring Using LLMs
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
Towards Verified Code Reasoning by LLMs
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
Bootstrapping Coding Agents: The Specification Is the Program
von: Monperrus, Martin
Veröffentlicht: (2026)
von: Monperrus, Martin
Veröffentlicht: (2026)
Grounding Data Science Code Generation with Input-Output Specifications
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
von: Wen, Yeming, et al.
Veröffentlicht: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
It's LIT! Reliability-Optimized LLMs with Inspectable Tools
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
Fast, Fine-Grained Equivalence Checking for Neural Decompilers
von: Dramko, Luke, et al.
Veröffentlicht: (2025)
von: Dramko, Luke, et al.
Veröffentlicht: (2025)
Automating API Documentation with LLMs: A BERTopic Approach
von: Naghshzan, AmirHossein
Veröffentlicht: (2025)
von: Naghshzan, AmirHossein
Veröffentlicht: (2025)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
iServe: An Intent-based Serving System for LLMs
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Liakopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
Breaking the Silence: the Threats of Using LLMs in Software Engineering
von: Sallou, June, et al.
Veröffentlicht: (2023)
von: Sallou, June, et al.
Veröffentlicht: (2023)
RepairBench: Leaderboard of Frontier Models for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2024)
von: Silva, André, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
Context-Augmented Code Generation Using Programming Knowledge Graphs
von: Seddik, Shahd, et al.
Veröffentlicht: (2026)
von: Seddik, Shahd, et al.
Veröffentlicht: (2026)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
Unravelling Technical debt topics through Time, Programming Languages and Repository
von: Shivashankar, Karthik, et al.
Veröffentlicht: (2025)
von: Shivashankar, Karthik, et al.
Veröffentlicht: (2025)
Programming with Pixels: Can Computer-Use Agents do Software Engineering?
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
von: Aggarwal, Pranjal, et al.
Veröffentlicht: (2025)
Automated Program Repair: Emerging trends pose and expose problems for benchmarks
von: Renzullo, Joseph, et al.
Veröffentlicht: (2024)
von: Renzullo, Joseph, et al.
Veröffentlicht: (2024)
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
von: Liu, Kaibo, et al.
Veröffentlicht: (2024)
von: Liu, Kaibo, et al.
Veröffentlicht: (2024)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
SkVM: Revisiting Language VM for Skills across Heterogenous LLMs and Harnesses
von: Chen, Le, et al.
Veröffentlicht: (2026)
von: Chen, Le, et al.
Veröffentlicht: (2026)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
Design, Implementation and Evaluation of a Novel Programming Language Topic Classification Workflow
von: Zhang, Michael, et al.
Veröffentlicht: (2025)
von: Zhang, Michael, et al.
Veröffentlicht: (2025)
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2023)
von: Silva, André, et al.
Veröffentlicht: (2023)
Beyond the Comfort Zone: Emerging Solutions to Overcome Challenges in Integrating LLMs into Software Products
von: Nahar, Nadia, et al.
Veröffentlicht: (2024)
von: Nahar, Nadia, et al.
Veröffentlicht: (2024)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2026)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
von: Errica, Federico, et al.
Veröffentlicht: (2024)
von: Errica, Federico, et al.
Veröffentlicht: (2024)
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics
von: Tran, Khang, et al.
Veröffentlicht: (2026)
von: Tran, Khang, et al.
Veröffentlicht: (2026)
ENCORE: Ensemble Learning using Convolution Neural Machine Translation for Automatic Program Repair
von: Lutellier, Thibaud, et al.
Veröffentlicht: (2019)
von: Lutellier, Thibaud, et al.
Veröffentlicht: (2019)
Ähnliche Einträge
-
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024) -
NExT: Teaching Large Language Models to Reason about Code Execution
von: Ni, Ansong, et al.
Veröffentlicht: (2024) -
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
von: Hooda, Ashish, et al.
Veröffentlicht: (2024) -
CigaR: Cost-efficient Program Repair with LLMs
von: Hidvégi, Dávid, et al.
Veröffentlicht: (2024) -
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)