Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yunseo, Song, John Youngeun, Kim, Dongsun, Kim, Jindae, Kim, Mijung, Nam, Jaechang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying Root Causes of Null Pointer Exceptions with Logical Inferences
by: Kim, Jindae, et al.
Published: (2024)
by: Kim, Jindae, et al.
Published: (2024)
Migrating Code At Scale With LLMs At Google
by: Ziftci, Celal, et al.
Published: (2025)
by: Ziftci, Celal, et al.
Published: (2025)
ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?
by: Nam, Doha, et al.
Published: (2025)
by: Nam, Doha, et al.
Published: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025)
by: Kim, Myeongsoo, et al.
Published: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
Code Hallucination
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
by: Rahman, Mirza Masfiqur, et al.
Published: (2024)
The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents
by: Kim, Yelin
Published: (2026)
by: Kim, Yelin
Published: (2026)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
A Low-Code Methodology for Developing AI Kiosks: a Case Study with the DIZEST Platform
by: Moon, SunMin, et al.
Published: (2025)
by: Moon, SunMin, et al.
Published: (2025)
Hallucination in LLM-Based Code Generation: An Automotive Case Study
by: Pavel, Marc, et al.
Published: (2025)
by: Pavel, Marc, et al.
Published: (2025)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
A Survey on Large Language Models for Code Generation
by: Jiang, Juyong, et al.
Published: (2024)
by: Jiang, Juyong, et al.
Published: (2024)
Eliminating Hallucination-Induced Errors in LLM Code Generation with Functional Clustering
by: Ravuri, Chaitanya, et al.
Published: (2025)
by: Ravuri, Chaitanya, et al.
Published: (2025)
Eliciting Instruction-tuned Code Language Models' Capabilities to Utilize Auxiliary Function for Code Generation
by: Lee, Seonghyeon, et al.
Published: (2024)
by: Lee, Seonghyeon, et al.
Published: (2024)
SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents
by: Lin, Feng, et al.
Published: (2024)
by: Lin, Feng, et al.
Published: (2024)
IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
by: Nekrasov, Roman, et al.
Published: (2025)
by: Nekrasov, Roman, et al.
Published: (2025)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
by: Han, Hojae, et al.
Published: (2024)
by: Han, Hojae, et al.
Published: (2024)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
by: Kim, Su-Hyeon, et al.
Published: (2025)
by: Kim, Su-Hyeon, et al.
Published: (2025)
Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
by: Liu, Chunhua, et al.
Published: (2025)
by: Liu, Chunhua, et al.
Published: (2025)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
by: Khati, Dipin, et al.
Published: (2026)
by: Khati, Dipin, et al.
Published: (2026)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
by: Ashraf, Humza, et al.
Published: (2025)
by: Ashraf, Humza, et al.
Published: (2025)
On the Impacts of Contexts on Repository-Level Code Generation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
by: Yoo, Jaeseok, et al.
Published: (2024)
by: Yoo, Jaeseok, et al.
Published: (2024)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
by: Islam, Md Arman, et al.
Published: (2025)
by: Islam, Md Arman, et al.
Published: (2025)
CodeMirage: Hallucinations in Code Generated by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
by: Hyun, Sangwon, et al.
Published: (2025)
by: Hyun, Sangwon, et al.
Published: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Benchmarking Correctness and Security in Multi-Turn Code Generation
by: Rawal, Ruchit, et al.
Published: (2025)
by: Rawal, Ruchit, et al.
Published: (2025)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
Lyra: A Benchmark for Turducken-Style Code Generation
by: Liang, Qingyuan, et al.
Published: (2021)
by: Liang, Qingyuan, et al.
Published: (2021)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
by: Xie, Danning, et al.
Published: (2025)
by: Xie, Danning, et al.
Published: (2025)
Similar Items
-
Identifying Root Causes of Null Pointer Exceptions with Logical Inferences
by: Kim, Jindae, et al.
Published: (2024) -
Migrating Code At Scale With LLMs At Google
by: Ziftci, Celal, et al.
Published: (2025) -
ReDef: Do Code Language Models Truly Understand Code Changes for Just-in-Time Software Defect Prediction?
by: Nam, Doha, et al.
Published: (2025) -
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024) -
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025)