Memorize or Generalize? Evaluating LLM Code Generation with Code Rewriting
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Lizhe, Chen, Wentao, Zhong, Li, Peng, Letian, Wang, Zilong, Shang, Jingbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Tale of LLMs and Induced Small Proxies: Scalable Agents for Knowledge Mining
by: Zhang, Sipeng, et al.
Published: (2025)
by: Zhang, Sipeng, et al.
Published: (2025)
The Price of Format: Diversity Collapse in LLMs
by: Yun, Longfei, et al.
Published: (2025)
by: Yun, Longfei, et al.
Published: (2025)
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
Training Language Models to Generate Quality Code with Program Analysis Feedback
by: Yao, Feng, et al.
Published: (2025)
by: Yao, Feng, et al.
Published: (2025)
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
by: Zhong, Li, et al.
Published: (2024)
by: Zhong, Li, et al.
Published: (2024)
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation
by: Zhong, Li, et al.
Published: (2023)
by: Zhong, Li, et al.
Published: (2023)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
by: Liu, Mingwei, et al.
Published: (2025)
by: Liu, Mingwei, et al.
Published: (2025)
Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
by: Zhao, Yicong, et al.
Published: (2025)
by: Zhao, Yicong, et al.
Published: (2025)
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
by: Fujii, Kazuki, et al.
Published: (2025)
by: Fujii, Kazuki, et al.
Published: (2025)
CodeFort: Robust Training for Code Generation Models
by: Zhang, Yuhao, et al.
Published: (2024)
by: Zhang, Yuhao, et al.
Published: (2024)
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
by: Peng, Wenjun, et al.
Published: (2026)
by: Peng, Wenjun, et al.
Published: (2026)
Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
by: Li, Jialin, et al.
Published: (2025)
by: Li, Jialin, et al.
Published: (2025)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
by: Hu, Wenhao, et al.
Published: (2025)
by: Hu, Wenhao, et al.
Published: (2025)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
by: Kim, Jungin, et al.
Published: (2025)
by: Kim, Jungin, et al.
Published: (2025)
An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc
by: Zhang, Hong, et al.
Published: (2026)
by: Zhang, Hong, et al.
Published: (2026)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Improving LLM-Generated Code Quality with GRPO
by: Robeyns, Maxime, et al.
Published: (2025)
by: Robeyns, Maxime, et al.
Published: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
Investigating The Smells of LLM Generated Code
by: Paul, Debalina Ghosh, et al.
Published: (2025)
by: Paul, Debalina Ghosh, et al.
Published: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot
by: Zhang, Beiqi, et al.
Published: (2024)
by: Zhang, Beiqi, et al.
Published: (2024)
Executing as You Generate: Hiding Execution Latency in LLM Code Generation
by: Sun, Zhensu, et al.
Published: (2026)
by: Sun, Zhensu, et al.
Published: (2026)
Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
by: Liu, Fang, et al.
Published: (2024)
by: Liu, Fang, et al.
Published: (2024)
MCCoder: Streamlining Motion Control with LLM-Assisted Code Generation and Rigorous Verification
by: Li, Yin, et al.
Published: (2024)
by: Li, Yin, et al.
Published: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
by: Wang, Jingquan, et al.
Published: (2025)
by: Wang, Jingquan, et al.
Published: (2025)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
by: Syromiatnikov, Mykyta, et al.
Published: (2025)
SAGE: Strategy-Adaptive Generation Engine for Query Rewriting
by: Wang, Teng, et al.
Published: (2025)
by: Wang, Teng, et al.
Published: (2025)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
by: Lian, Keke, et al.
Published: (2025)
by: Lian, Keke, et al.
Published: (2025)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
by: Wang, Xinchen, et al.
Published: (2025)
by: Wang, Xinchen, et al.
Published: (2025)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's Nest
by: Peng, Letian, et al.
Published: (2025)
by: Peng, Letian, et al.
Published: (2025)
Reasoning-Driven Multimodal LLM for Domain Generalization
by: Xu, Zhipeng, et al.
Published: (2026)
by: Xu, Zhipeng, et al.
Published: (2026)
Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
Evaluating SAP Joule for Code Generation
by: Heisler, Joshua, et al.
Published: (2025)
by: Heisler, Joshua, et al.
Published: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
by: Peng, Yun, et al.
Published: (2024)
by: Peng, Yun, et al.
Published: (2024)
Similar Items
-
A Tale of LLMs and Induced Small Proxies: Scalable Agents for Knowledge Mining
by: Zhang, Sipeng, et al.
Published: (2025) -
The Price of Format: Diversity Collapse in LLMs
by: Yun, Longfei, et al.
Published: (2025) -
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
by: Ye, Tong, et al.
Published: (2024) -
Training Language Models to Generate Quality Code with Program Analysis Feedback
by: Yao, Feng, et al.
Published: (2025) -
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
by: Zhong, Li, et al.
Published: (2024)