Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
Fuente:
arXiv
Saved in:
| Main Authors: | He, Kaifeng, Zhang, Xiaojun, Cai, Peiliang, Liu, Mingwei, Wang, Yanlin, Wang, Chong, Huang, Kaifeng, Chen, Bihuan, Peng, Xin, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Preliminary Study on the Robustness of Code Generation by Large Language Models
by: Li, Zike, et al.
Published: (2025)
by: Li, Zike, et al.
Published: (2025)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025)
by: He, Kaifeng, et al.
Published: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
by: Duan, Guoliang, et al.
Published: (2025)
by: Duan, Guoliang, et al.
Published: (2025)
Scalable and Precise Application-Centered Call Graph Construction for Python
by: Huang, Kaifeng, et al.
Published: (2023)
by: Huang, Kaifeng, et al.
Published: (2023)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
by: Ou, Guangsheng, et al.
Published: (2024)
by: Ou, Guangsheng, et al.
Published: (2024)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
Killing Two Birds with One Stone: Malicious Package Detection in NPM and PyPI using a Single Model of Malicious Behavior Sequence
by: Zhang, Junan, et al.
Published: (2023)
by: Zhang, Junan, et al.
Published: (2023)
TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
by: Lu, You, et al.
Published: (2025)
by: Lu, You, et al.
Published: (2025)
LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Generating High-Quality Datasets for Code Editing via Open-Source Language Models
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
by: Liu, Mingwei, et al.
Published: (2026)
by: Liu, Mingwei, et al.
Published: (2026)
DRAINCODE: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context Poisoning
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
Evolving Triple Knowledge-Augmented LLMs for Code Translation in Repository Context
by: Ou, Guangsheng, et al.
Published: (2025)
by: Ou, Guangsheng, et al.
Published: (2025)
Identifying Smart Contract Security Issues in Code Snippets from Stack Overflow
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study
by: Liu, Mingwei, et al.
Published: (2025)
by: Liu, Mingwei, et al.
Published: (2025)
Dynamic analysis enhances issue resolution
by: Liu, Mingwei, et al.
Published: (2026)
by: Liu, Mingwei, et al.
Published: (2026)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
by: Dai, Dekun, et al.
Published: (2025)
by: Dai, Dekun, et al.
Published: (2025)
What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
by: Gu, Wenchao, et al.
Published: (2025)
by: Gu, Wenchao, et al.
Published: (2025)
EffiReasonTrans: RL-Optimized Reasoning for Code Translation
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Are Decoder-Only Large Language Models the Silver Bullet for Code Search?
by: Chen, Yuxuan, et al.
Published: (2024)
by: Chen, Yuxuan, et al.
Published: (2024)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
Code Digital Twin: Empowering LLMs with Tacit Knowledge for Complex Software Development
by: Peng, Xin, et al.
Published: (2025)
by: Peng, Xin, et al.
Published: (2025)
Lifting the Veil on Composition, Risks, and Mitigations of the Large Language Model Supply Chain
by: Huang, Kaifeng, et al.
Published: (2024)
by: Huang, Kaifeng, et al.
Published: (2024)
KTester: Leveraging Domain and Testing Knowledge for More Effective LLM-based Test Generation
by: Li, Anji, et al.
Published: (2025)
by: Li, Anji, et al.
Published: (2025)
A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
by: Gao, Cuiyun, et al.
Published: (2025)
by: Gao, Cuiyun, et al.
Published: (2025)
SparseCoder: Identifier-Aware Sparse Transformer for File-Level Code Summarization
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
Build-Aware Incremental C-to-Rust Migration via Skeleton-First Translation and Historical Knowledge Reuse
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
RepoDoc: A Knowledge Graph-Based Framework to Automatic Documentation Generation and Incremental Updates
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
RLCoder: Reinforcement Learning for Repository-Level Code Completion
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
by: Zheng, Zibin, et al.
Published: (2023)
by: Zheng, Zibin, et al.
Published: (2023)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
DynNPC: Finding More Violations Induced by ADS in Simulation Testing through Dynamic NPC Behavior Generation
by: Lu, You, et al.
Published: (2024)
by: Lu, You, et al.
Published: (2024)
AnomalyGen: Enhancing Log-Based Anomaly Detection with Code-Guided Data Augmentation
by: Li, Xinyu, et al.
Published: (2026)
by: Li, Xinyu, et al.
Published: (2026)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
RoadGen: Generating Road Scenarios for Autonomous Vehicle Testing
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
Code Digital Twin: A Knowledge Infrastructure for AI-Assisted Complex Software Development
by: Peng, Xin, et al.
Published: (2025)
by: Peng, Xin, et al.
Published: (2025)
Similar Items
-
A Preliminary Study on the Robustness of Code Generation by Large Language Models
by: Li, Zike, et al.
Published: (2025) -
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
by: He, Kaifeng, et al.
Published: (2025) -
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
by: Duan, Guoliang, et al.
Published: (2025) -
Scalable and Precise Application-Centered Call Graph Construction for Python
by: Huang, Kaifeng, et al.
Published: (2023) -
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
by: Ou, Guangsheng, et al.
Published: (2024)