Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ngassom, Sylvain Kouemo, Dakhel, Arghavan Moradi, Tambon, Florian, Khomh, Foutse |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024)
by: Da Silva, Leuson, et al.
Published: (2024)
GIST: Generated Inputs Sets Transferability in Deep Learning
by: Tambon, Florian, et al.
Published: (2023)
by: Tambon, Florian, et al.
Published: (2023)
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
by: Majdinasab, Vahid, et al.
Published: (2025)
by: Majdinasab, Vahid, et al.
Published: (2025)
Inferring Code Correctness from Specification
by: Florian, Tambon, et al.
Published: (2026)
by: Florian, Tambon, et al.
Published: (2026)
ReCatcher: Towards LLMs Regression Testing for Code Generation
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
by: Abukhalaf, Seif, et al.
Published: (2024)
by: Abukhalaf, Seif, et al.
Published: (2024)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
by: Oueslati, Khouloud, et al.
Published: (2025)
by: Oueslati, Khouloud, et al.
Published: (2025)
Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
by: Jamshidi, Saeid, et al.
Published: (2025)
by: Jamshidi, Saeid, et al.
Published: (2025)
Machine Learning Robustness: A Primer
by: Braiek, Houssem Ben, et al.
Published: (2024)
by: Braiek, Houssem Ben, et al.
Published: (2024)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
by: Morovati, Mohammad Mehdi, et al.
Published: (2023)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
by: Aghili, Roozbeh, et al.
Published: (2025)
by: Aghili, Roozbeh, et al.
Published: (2025)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
by: Wu, Xingfang, et al.
Published: (2024)
by: Wu, Xingfang, et al.
Published: (2024)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
by: Taraghi, Mina, et al.
Published: (2024)
by: Taraghi, Mina, et al.
Published: (2024)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
by: Shah, Mehil B, et al.
Published: (2025)
by: Shah, Mehil B, et al.
Published: (2025)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
by: Jamshidi, Saeid, et al.
Published: (2025)
by: Jamshidi, Saeid, et al.
Published: (2025)
Are LLMs Reliable Code Reviewers? Systematic Overcorrection in Requirement Conformance Judgement
by: Jin, Haolin, et al.
Published: (2026)
by: Jin, Haolin, et al.
Published: (2026)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
by: Dolcetti, Greta, et al.
Published: (2024)
by: Dolcetti, Greta, et al.
Published: (2024)
$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware Selective Contrastive Decoding
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
by: Cramer, Marcos, et al.
Published: (2025)
by: Cramer, Marcos, et al.
Published: (2025)
A Taxonomy of Inefficiencies in LLM-Generated Python Code
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
Improving the Robustness of Large Language Models for Code Tasks via Fine-tuning with Perturbed Data
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
by: Bouchoucha, Rached, et al.
Published: (2024)
by: Bouchoucha, Rached, et al.
Published: (2024)
On the Effectiveness of LLMs for Manual Test Verifications
by: Peixoto, Myron David Lucena Campos, et al.
Published: (2024)
by: Peixoto, Myron David Lucena Campos, et al.
Published: (2024)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
by: Islam, Md Arman, et al.
Published: (2025)
by: Islam, Md Arman, et al.
Published: (2025)
LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
Adversarial Moral Stress Testing of Large Language Models
by: Jamshidi, Saeid, et al.
Published: (2026)
by: Jamshidi, Saeid, et al.
Published: (2026)
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
FastCoder: Accelerating Repository-level Code Generation via Efficient Retrieval and Verification
by: Zhao, Qianhui, et al.
Published: (2025)
by: Zhao, Qianhui, et al.
Published: (2025)
From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification
by: Erfan, Md, et al.
Published: (2026)
by: Erfan, Md, et al.
Published: (2026)
From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation
by: Yang, Guang, et al.
Published: (2026)
by: Yang, Guang, et al.
Published: (2026)
RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation
by: Zhong, Sicheng, et al.
Published: (2025)
by: Zhong, Sicheng, et al.
Published: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Similar Items
-
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
by: Tambon, Florian, et al.
Published: (2024) -
Refactoring with LLMs: Bridging Human Expertise and Machine Understanding
by: Piao, Yonnel Chen Kuang, et al.
Published: (2025) -
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
by: Da Silva, Leuson, et al.
Published: (2024) -
GIST: Generated Inputs Sets Transferability in Deep Learning
by: Tambon, Florian, et al.
Published: (2023)