Towards Verified Code Reasoning by LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sistla, Meghana, Balakrishnan, Gogul, Rondon, Pat, Cambronero, José, Tufano, Michele, Chandra, Satish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Agent-based Program Repair at Google
von: Rondon, Pat, et al.
Veröffentlicht: (2025)
von: Rondon, Pat, et al.
Veröffentlicht: (2025)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
von: Cheng, Runxiang, et al.
Veröffentlicht: (2025)
von: Cheng, Runxiang, et al.
Veröffentlicht: (2025)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
von: Shi, Sherry, et al.
Veröffentlicht: (2025)
von: Shi, Sherry, et al.
Veröffentlicht: (2025)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
von: Khati, Dipin, et al.
Veröffentlicht: (2025)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
von: Steenhoek, Benjamin, et al.
Veröffentlicht: (2023)
von: Steenhoek, Benjamin, et al.
Veröffentlicht: (2023)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
von: Steenhoek, Benjamin, et al.
Veröffentlicht: (2024)
von: Steenhoek, Benjamin, et al.
Veröffentlicht: (2024)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
von: Cheng, Runxiang, et al.
Veröffentlicht: (2026)
von: Cheng, Runxiang, et al.
Veröffentlicht: (2026)
Leveraging Reward Models for Guiding Code Review Comment Generation
von: Sghaier, Oussama Ben, et al.
Veröffentlicht: (2025)
von: Sghaier, Oussama Ben, et al.
Veröffentlicht: (2025)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
von: Cambronero, José, et al.
Veröffentlicht: (2025)
von: Cambronero, José, et al.
Veröffentlicht: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
von: Haque, Mirazul, et al.
Veröffentlicht: (2025)
Verifier-Guided Code Translation via Meta-Step Decoding
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
CRQBench: A Benchmark of Code Reasoning Questions
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2024)
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2024)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
Agentic Code Reasoning
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
Teaching Code Refactoring Using LLMs
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
von: Khairnar, Anshul, et al.
Veröffentlicht: (2025)
Clover: Closed-Loop Verifiable Code Generation
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
von: Sun, Chuyue, et al.
Veröffentlicht: (2023)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
von: Palacio, David N., et al.
Veröffentlicht: (2024)
von: Palacio, David N., et al.
Veröffentlicht: (2024)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
von: Diggs, Colin, et al.
Veröffentlicht: (2024)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
von: Xie, Zichen, et al.
Veröffentlicht: (2026)
CONCORD: Towards a DSL for Configurable Graph Code Representation
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
von: Saad, Mootez, et al.
Veröffentlicht: (2024)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
von: Miserendino, Samuel, et al.
Veröffentlicht: (2025)
von: Miserendino, Samuel, et al.
Veröffentlicht: (2025)
Automating Code Review: A Systematic Literature Review
von: Tufano, Rosalia, et al.
Veröffentlicht: (2025)
von: Tufano, Rosalia, et al.
Veröffentlicht: (2025)
Verifying Machine Learning Interpretability Requirements through Provenance
von: Vonderhaar, Lynn, et al.
Veröffentlicht: (2026)
von: Vonderhaar, Lynn, et al.
Veröffentlicht: (2026)
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
von: Mhatre, Akshay, et al.
Veröffentlicht: (2025)
von: Mhatre, Akshay, et al.
Veröffentlicht: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
von: Zhang, Wuyang, et al.
Veröffentlicht: (2024)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
von: Nazzal, Mahmoud, et al.
Veröffentlicht: (2024)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
On LLMs' Internal Representation of Code Correctness
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
von: Ribeiro, Francisco, et al.
Veröffentlicht: (2025)
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating Agent-based Program Repair at Google
von: Rondon, Pat, et al.
Veröffentlicht: (2025) -
Agentic Bug Reproduction for Effective Automated Program Repair at Google
von: Cheng, Runxiang, et al.
Veröffentlicht: (2025) -
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
von: Shi, Sherry, et al.
Veröffentlicht: (2025) -
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
von: Khati, Dipin, et al.
Veröffentlicht: (2025) -
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)