CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yuchen, Yan, Weixiang, Yang, Qian, Zhao, Xuandong, Chen, Qian, Wang, Wen, Luo, Ziyang, Ma, Lei, Song, Dawn |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
by: Zhao, Yuwei, et al.
Published: (2024)
by: Zhao, Yuwei, et al.
Published: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026)
by: Shi, Jieke, et al.
Published: (2026)
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
by: Chen, Yujia, et al.
Published: (2025)
by: Chen, Yujia, et al.
Published: (2025)
CodeScore: Evaluating Code Generation by Learning Code Execution
by: Dong, Yihong, et al.
Published: (2023)
by: Dong, Yihong, et al.
Published: (2023)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
SelfPiCo: Self-Guided Partial Code Execution with LLMs
by: Xue, Zhipeng, et al.
Published: (2024)
by: Xue, Zhipeng, et al.
Published: (2024)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
by: Wang, Junqiao, et al.
Published: (2024)
by: Wang, Junqiao, et al.
Published: (2024)
Code Hallucination
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
Python Symbolic Execution with LLM-powered Code Generation
by: Wang, Wenhan, et al.
Published: (2024)
by: Wang, Wenhan, et al.
Published: (2024)
Dissect-and-Restore: AI-based Code Verification with Transient Refactoring
by: Wang, Changjie, et al.
Published: (2025)
by: Wang, Changjie, et al.
Published: (2025)
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
by: Sakharova, Marina, et al.
Published: (2025)
by: Sakharova, Marina, et al.
Published: (2025)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
by: Li, Kaixin, et al.
Published: (2024)
by: Li, Kaixin, et al.
Published: (2024)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
by: Zheng, Tianyu, et al.
Published: (2024)
by: Zheng, Tianyu, et al.
Published: (2024)
Integrated Modeling, Verification, and Code Generation for Unmanned Aerial Systems
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
A Benchmark for Localizing Code and Non-Code Issues in Software Projects
by: Zhang, Zejun, et al.
Published: (2025)
by: Zhang, Zejun, et al.
Published: (2025)
Issue Retrieval and Verification Enhanced Supplementary Code Comment Generation
by: Zou, Yanzhen, et al.
Published: (2025)
by: Zou, Yanzhen, et al.
Published: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
by: Guo, Chengquan, et al.
Published: (2025)
by: Guo, Chengquan, et al.
Published: (2025)
On Evaluating the Efficiency of Source Code Generated by LLMs
by: Niu, Changan, et al.
Published: (2024)
by: Niu, Changan, et al.
Published: (2024)
Beyond Execution: Static-Analysis Rewards and Hint-Conditioned Diffusion RL for Code Generation
by: Ouyang, Shuyin, et al.
Published: (2026)
by: Ouyang, Shuyin, et al.
Published: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
by: Haque, Mirazul, et al.
Published: (2025)
by: Haque, Mirazul, et al.
Published: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
by: Eghbali, Aryaz, et al.
Published: (2024)
by: Eghbali, Aryaz, et al.
Published: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
Hybrid-Code v2: Zero-Hallucination Clinical ICD-10 Coding via Neuro-Symbolic Verification and Automated Knowledge Base Expansion
by: Yu, Yunguo
Published: (2025)
by: Yu, Yunguo
Published: (2025)
RepoZero: Can LLMs Generate a Code Repository from Scratch?
by: Zhang, Zhaoxi, et al.
Published: (2026)
by: Zhang, Zhaoxi, et al.
Published: (2026)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
by: Dai, Yihan, et al.
Published: (2025)
by: Dai, Yihan, et al.
Published: (2025)
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
by: Li, Zenan, et al.
Published: (2026)
by: Li, Zenan, et al.
Published: (2026)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
by: He, Xinyi, et al.
Published: (2025)
by: He, Xinyi, et al.
Published: (2025)
ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
by: He, Minghua, et al.
Published: (2025)
by: He, Minghua, et al.
Published: (2025)
Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
Aligning Requirement for Large Language Model's Code Generation
by: Tian, Zhao, et al.
Published: (2025)
by: Tian, Zhao, et al.
Published: (2025)
Unmasking the Genuine Type Inference Capabilities of LLMs for Java Code Snippets
by: Dong, Yiwen, et al.
Published: (2025)
by: Dong, Yiwen, et al.
Published: (2025)
A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
by: Gao, Cuiyun, et al.
Published: (2025)
by: Gao, Cuiyun, et al.
Published: (2025)
Similar Items
-
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
by: Zhao, Yuwei, et al.
Published: (2024) -
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023) -
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024) -
Compiling Code LLMs into Lightweight Executables
by: Shi, Jieke, et al.
Published: (2026) -
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
by: Chen, Yujia, et al.
Published: (2025)