A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Cuiyun, Fan, Guodong, Chong, Chun Yong, Chen, Shizhan, Liu, Chao, Lo, David, Zheng, Zibin, Liao, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Model Editing Meets Service Evolution: A Knowledge-Update Perspective for Service Recommendation
by: Fan, Guodong, et al.
Published: (2026)
by: Fan, Guodong, et al.
Published: (2026)
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
by: Chen, Yujia, et al.
Published: (2025)
by: Chen, Yujia, et al.
Published: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Deep Learning Based Code Generation Methods: Literature Review
by: Yang, Zezhou, et al.
Published: (2023)
by: Yang, Zezhou, et al.
Published: (2023)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
by: Feng, Jia, et al.
Published: (2024)
by: Feng, Jia, et al.
Published: (2024)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
by: He, Kaifeng, et al.
Published: (2026)
by: He, Kaifeng, et al.
Published: (2026)
APIGen: Generative API Method Recommendation
by: Chen, Yujia, et al.
Published: (2024)
by: Chen, Yujia, et al.
Published: (2024)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
A Roadmap on Modern Code Review: Challenges and Opportunities
by: Yang, Zezhou, et al.
Published: (2024)
by: Yang, Zezhou, et al.
Published: (2024)
A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
by: Li, Yaoxian, et al.
Published: (2022)
by: Li, Yaoxian, et al.
Published: (2022)
Bridge and Hint: Extending Pre-trained Language Models for Long-Range Code
by: Chen, Yujia, et al.
Published: (2024)
by: Chen, Yujia, et al.
Published: (2024)
Are LLMs Reliable Code Reviewers? Systematic Overcorrection in Requirement Conformance Judgement
by: Jin, Haolin, et al.
Published: (2026)
by: Jin, Haolin, et al.
Published: (2026)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
Automating Code Review: A Systematic Literature Review
by: Tufano, Rosalia, et al.
Published: (2025)
by: Tufano, Rosalia, et al.
Published: (2025)
Vibe Coding in Practice: Motivations, Challenges, and a Future Outlook -- a Grey Literature Review
by: Fawzy, Ahmed, et al.
Published: (2025)
by: Fawzy, Ahmed, et al.
Published: (2025)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
by: Zhan, Zexun, et al.
Published: (2025)
by: Zhan, Zexun, et al.
Published: (2025)
A Systematic Literature Review on Neural Code Translation
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Language Models for Code Optimization: Survey, Challenges and Future Directions
by: Gong, Jingzhi, et al.
Published: (2025)
by: Gong, Jingzhi, et al.
Published: (2025)
Search-Based LLMs for Code Optimization
by: Gao, Shuzheng, et al.
Published: (2024)
by: Gao, Shuzheng, et al.
Published: (2024)
MLLM-Based UI2Code Automation Guided by UI Layout Information
by: Wu, Fan, et al.
Published: (2025)
by: Wu, Fan, et al.
Published: (2025)
Software Architecture Meets LLMs: A Systematic Literature Review
by: Schmid, Larissa, et al.
Published: (2025)
by: Schmid, Larissa, et al.
Published: (2025)
An Empirical Analysis of Static Analysis Methods for Detection and Mitigation of Code Library Hallucinations
by: Miranda-Pena, Clarissa, et al.
Published: (2026)
by: Miranda-Pena, Clarissa, et al.
Published: (2026)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
by: Eghbali, Aryaz, et al.
Published: (2024)
by: Eghbali, Aryaz, et al.
Published: (2024)
LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda
by: Gomes, Victoria, et al.
Published: (2026)
by: Gomes, Victoria, et al.
Published: (2026)
Understanding Web Application Workloads and Their Applications: Systematic Literature Review and Characterization
by: Aghili, Roozbeh, et al.
Published: (2024)
by: Aghili, Roozbeh, et al.
Published: (2024)
Android Source Code Smells: A Systematic Literature Review
by: Muhammad Fawad, et al.
Published: (2024)
by: Muhammad Fawad, et al.
Published: (2024)
Maintainability Challenges in ML: A Systematic Literature Review
by: Shivashankar, Karthik, et al.
Published: (2024)
by: Shivashankar, Karthik, et al.
Published: (2024)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
by: Yang, Zezhou, et al.
Published: (2025)
by: Yang, Zezhou, et al.
Published: (2025)
Can LLMs be Effective Code Contributors? A Study on Open-source Projects
by: Chong, Chun Jie, et al.
Published: (2026)
by: Chong, Chun Jie, et al.
Published: (2026)
Security of Language Models for Code: A Systematic Literature Review
by: Chen, Yuchen, et al.
Published: (2024)
by: Chen, Yuchen, et al.
Published: (2024)
Prompt-Driven Code Summarization: A Systematic Literature Review
by: Farjana, Afia, et al.
Published: (2026)
by: Farjana, Afia, et al.
Published: (2026)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
by: Gao, Shuzheng, et al.
Published: (2023)
by: Gao, Shuzheng, et al.
Published: (2023)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026)
by: Hu, Ruida, et al.
Published: (2026)
Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
by: Chen, Yujia, et al.
Published: (2026)
by: Chen, Yujia, et al.
Published: (2026)
A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions
by: Chochlov, Muslim, et al.
Published: (2026)
by: Chochlov, Muslim, et al.
Published: (2026)
Similar Items
-
When Model Editing Meets Service Evolution: A Knowledge-Update Perspective for Service Recommendation
by: Fan, Guodong, et al.
Published: (2026) -
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
by: Chen, Yujia, et al.
Published: (2025) -
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024) -
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025) -
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)