LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Ziyao, Wang, Yanlin, Wang, Chong, Chen, Jiachi, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Towards an Understanding of Context Utilization in Code Intelligence
par: Wang, Yanlin, et autres
Publié: (2025)
par: Wang, Yanlin, et autres
Publié: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
par: Zheng, Dewu, et autres
Publié: (2024)
par: Zheng, Dewu, et autres
Publié: (2024)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
par: Wang, Yanli, et autres
Publié: (2024)
par: Wang, Yanli, et autres
Publié: (2024)
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
par: Liu, Sicong, et autres
Publié: (2026)
par: Liu, Sicong, et autres
Publié: (2026)
Agents in Software Engineering: Survey, Landscape, and Vision
par: Wang, Yanlin, et autres
Publié: (2024)
par: Wang, Yanlin, et autres
Publié: (2024)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
par: Wang, Yanlin, et autres
Publié: (2025)
par: Wang, Yanlin, et autres
Publié: (2025)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
par: Jiang, Tianyue, et autres
Publié: (2026)
par: Jiang, Tianyue, et autres
Publié: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
par: Liang, Linxi, et autres
Publié: (2025)
par: Liang, Linxi, et autres
Publié: (2025)
Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
par: Wang, Yanlin, et autres
Publié: (2024)
par: Wang, Yanlin, et autres
Publié: (2024)
CodeMirage: Hallucinations in Code Generated by Large Language Models
par: Agarwal, Vibhor, et autres
Publié: (2024)
par: Agarwal, Vibhor, et autres
Publié: (2024)
DevEval: Evaluating Code Generation in Practical Software Projects
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
par: Wang, Yanlin, et autres
Publié: (2026)
par: Wang, Yanlin, et autres
Publié: (2026)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
par: He, Kaifeng, et autres
Publié: (2026)
par: He, Kaifeng, et autres
Publié: (2026)
IndustryCode: A Benchmark for Industry Code Generation
par: Zeng, Puyu, et autres
Publié: (2026)
par: Zeng, Puyu, et autres
Publié: (2026)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
par: Zheng, Tianyu, et autres
Publié: (2024)
par: Zheng, Tianyu, et autres
Publié: (2024)
ETF: An Entity Tracing Framework for Hallucination Detection in Code Summaries
par: Maharaj, Kishan, et autres
Publié: (2024)
par: Maharaj, Kishan, et autres
Publié: (2024)
Pragmatic Reasoning improves LLM Code Generation
par: Cao, Zhuchen, et autres
Publié: (2025)
par: Cao, Zhuchen, et autres
Publié: (2025)
Mitigating Gender Bias in Code Large Language Models via Model Editing
par: Qin, Zhanyue, et autres
Publié: (2024)
par: Qin, Zhanyue, et autres
Publié: (2024)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
par: Gao, Shuzheng, et autres
Publié: (2025)
par: Gao, Shuzheng, et autres
Publié: (2025)
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
par: Cao, Jialun, et autres
Publié: (2025)
par: Cao, Jialun, et autres
Publié: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
par: Guo, Lianghong, et autres
Publié: (2025)
par: Guo, Lianghong, et autres
Publié: (2025)
Magicoder: Empowering Code Generation with OSS-Instruct
par: Wei, Yuxiang, et autres
Publié: (2023)
par: Wei, Yuxiang, et autres
Publié: (2023)
LocAgent: Graph-Guided LLM Agents for Code Localization
par: Chen, Zhaoling, et autres
Publié: (2025)
par: Chen, Zhaoling, et autres
Publié: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
par: Tao, Tianhua, et autres
Publié: (2024)
par: Tao, Tianhua, et autres
Publié: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
par: Yan, Weixiang, et autres
Publié: (2023)
par: Yan, Weixiang, et autres
Publié: (2023)
Generating High-Quality Datasets for Code Editing via Open-Source Language Models
par: Zhang, Zekai, et autres
Publié: (2025)
par: Zhang, Zekai, et autres
Publié: (2025)
A Taxonomy of Prompt Defects in LLM Systems
par: Tian, Haoye, et autres
Publié: (2025)
par: Tian, Haoye, et autres
Publié: (2025)
A Survey on Code Generation with LLM-based Agents
par: Dong, Yihong, et autres
Publié: (2025)
par: Dong, Yihong, et autres
Publié: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
par: Peng, Yun, et autres
Publié: (2024)
par: Peng, Yun, et autres
Publié: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
par: Zhang, William, et autres
Publié: (2024)
par: Zhang, William, et autres
Publié: (2024)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
par: Zhang, Tianyi, et autres
Publié: (2025)
par: Zhang, Tianyi, et autres
Publié: (2025)
Multilingual Multimodal Software Developer for Code Generation
par: Chai, Linzheng, et autres
Publié: (2025)
par: Chai, Linzheng, et autres
Publié: (2025)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
par: Sapkota, Ranjan, et autres
Publié: (2025)
par: Sapkota, Ranjan, et autres
Publié: (2025)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
par: Shi, Yuling, et autres
Publié: (2024)
par: Shi, Yuling, et autres
Publié: (2024)
Identifying Smart Contract Security Issues in Code Snippets from Stack Overflow
par: Chen, Jiachi, et autres
Publié: (2024)
par: Chen, Jiachi, et autres
Publié: (2024)
Verification Limits Code LLM Training
par: Gureja, Srishti, et autres
Publié: (2025)
par: Gureja, Srishti, et autres
Publié: (2025)
A Survey on Large Language Models for Code Generation
par: Jiang, Juyong, et autres
Publié: (2024)
par: Jiang, Juyong, et autres
Publié: (2024)
Documents similaires
-
Towards an Understanding of Context Utilization in Code Intelligence
par: Wang, Yanlin, et autres
Publié: (2025) -
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
par: Zheng, Dewu, et autres
Publié: (2024) -
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
par: Wang, Yanli, et autres
Publié: (2024) -
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
par: Liu, Sicong, et autres
Publié: (2026) -
Agents in Software Engineering: Survey, Landscape, and Vision
par: Wang, Yanlin, et autres
Publié: (2024)