On the Effectiveness of Training Data Optimization for LLM-based Code Generation: An Empirical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Shiqi, Tian, Zhao, Xiao, Tao, Wang, Dong, Chen, Junjie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution
by: Kuang, Shiqi, et al.
Published: (2026)
by: Kuang, Shiqi, et al.
Published: (2026)
Aligning Requirement for Large Language Model's Code Generation
by: Tian, Zhao, et al.
Published: (2025)
by: Tian, Zhao, et al.
Published: (2025)
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
by: Dong, Shijia, et al.
Published: (2026)
by: Dong, Shijia, et al.
Published: (2026)
"Refactoring Runaway": Understanding and Mitigating Tangled Refactorings in Coding Agents for Issue Resolution
by: Tian, Zhao, et al.
Published: (2026)
by: Tian, Zhao, et al.
Published: (2026)
What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
by: Gu, Wenchao, et al.
Published: (2025)
by: Gu, Wenchao, et al.
Published: (2025)
Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
by: Tian, Zhao, et al.
Published: (2023)
by: Tian, Zhao, et al.
Published: (2023)
Improving LLM Code Generation via Requirement-Aware Curriculum Reinforcement Learning
by: Yin, Shouyu, et al.
Published: (2026)
by: Yin, Shouyu, et al.
Published: (2026)
An Empirical Study of Interaction Smells in Multi-Turn Human-LLM Collaborative Code Generation
by: Zhang, Binquan, et al.
Published: (2026)
by: Zhang, Binquan, et al.
Published: (2026)
Large Language Models for Mobile GUI Text Input Generation: An Empirical Study
by: Cui, Chenhui, et al.
Published: (2024)
by: Cui, Chenhui, et al.
Published: (2024)
Understanding Chain-of-Thought Effectiveness in Code Generation: An Empirical and Information-Theoretic Analysis
by: Jin, Naizhu, et al.
Published: (2025)
by: Jin, Naizhu, et al.
Published: (2025)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
by: Vartziotis, Tina, et al.
Published: (2024)
by: Vartziotis, Tina, et al.
Published: (2024)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
GrepRAG: An Empirical Study and Optimization of Grep-Like Retrieval for Code Completion
by: Wang, Baoyi, et al.
Published: (2026)
by: Wang, Baoyi, et al.
Published: (2026)
An Empirical Study of LLM-Based Code Clone Detection
by: Zhu, Wenqing, et al.
Published: (2025)
by: Zhu, Wenqing, et al.
Published: (2025)
LLM-Based Test-Driven Interactive Code Generation: User Study and Empirical Evaluation
by: Fakhoury, Sarah, et al.
Published: (2024)
by: Fakhoury, Sarah, et al.
Published: (2024)
Toward Effective Secure Code Reviews: An Empirical Study of Security-Related Coding Weaknesses
by: Charoenwet, Wachiraphan, et al.
Published: (2023)
by: Charoenwet, Wachiraphan, et al.
Published: (2023)
SemOpt: LLM-Driven Code Optimization via Rule-Based Analysis
by: Zhao, Yuwei, et al.
Published: (2025)
by: Zhao, Yuwei, et al.
Published: (2025)
Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
by: Aðalsteinsson, Fannar Steinn, et al.
Published: (2025)
by: Aðalsteinsson, Fannar Steinn, et al.
Published: (2025)
A Large-Scale Empirical Study of AI-Generated Code in Real-World Repositories
by: Mao, Tianhao, et al.
Published: (2026)
by: Mao, Tianhao, et al.
Published: (2026)
An Empirical Study of the Non-determinism of ChatGPT in Code Generation
by: Ouyang, Shuyin, et al.
Published: (2023)
by: Ouyang, Shuyin, et al.
Published: (2023)
Compact Constraint Encoding for LLM Code Generation: An Empirical Study of Token Economics and Constraint Compliance
by: Tang, Hanzhang
Published: (2026)
by: Tang, Hanzhang
Published: (2026)
Detect Repair Verify for Securing LLM Generated Code: A Multi-Language Empirical Study
by: Cheng, Cheng
Published: (2026)
by: Cheng, Cheng
Published: (2026)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
by: Yang, Zezhou, et al.
Published: (2025)
by: Yang, Zezhou, et al.
Published: (2025)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
An Empirical Study of Developers' Challenges in Implementing Workflows as Code: A Case Study on Apache Airflow
by: Yasmin, Jerin, et al.
Published: (2024)
by: Yasmin, Jerin, et al.
Published: (2024)
LLM-based Vulnerability Detection at Project Scale: An Empirical Study
by: Li, Fengjie, et al.
Published: (2026)
by: Li, Fengjie, et al.
Published: (2026)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
by: Crupi, Giuseppe, et al.
Published: (2025)
by: Crupi, Giuseppe, et al.
Published: (2025)
Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
Detect--Repair--Verify for LLM-Generated Code: A Multi-Language, Multi-Granularity Empirical Study
by: Cheng, Cheng
Published: (2026)
by: Cheng, Cheng
Published: (2026)
Failure-Aware Enhancements for Large Language Model (LLM) Code Generation: An Empirical Study on Decision Framework
by: Shen, Jianru, et al.
Published: (2026)
by: Shen, Jianru, et al.
Published: (2026)
Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
by: Zhang, Binquan, et al.
Published: (2025)
by: Zhang, Binquan, et al.
Published: (2025)
Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study
by: Li, Yichen, et al.
Published: (2023)
by: Li, Yichen, et al.
Published: (2023)
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
Code Review Automation Via Multi-task Federated LLM -- An Empirical Study
by: Kumar, Jahnavi, et al.
Published: (2024)
by: Kumar, Jahnavi, et al.
Published: (2024)
Are Coding Agents Generating Over-Mocked Tests? An Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
An Empirical Study on the Effectiveness of Large Language Models for Binary Code Understanding
by: Shang, Xiuwei, et al.
Published: (2025)
by: Shang, Xiuwei, et al.
Published: (2025)
Optimization-Aware Test Generation for Deep Learning Compilers
by: Shen, Qingchao, et al.
Published: (2025)
by: Shen, Qingchao, et al.
Published: (2025)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
by: Dong, Zeming, et al.
Published: (2023)
by: Dong, Zeming, et al.
Published: (2023)
Benchmarking and Studying the LLM-based Code Review
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
An Empirical Study on Challenges for LLM Application Developers
by: Chen, Xiang, et al.
Published: (2024)
by: Chen, Xiang, et al.
Published: (2024)
Similar Items
-
REAgent: Requirement-Driven LLM Agents for Software Issue Resolution
by: Kuang, Shiqi, et al.
Published: (2026) -
Aligning Requirement for Large Language Model's Code Generation
by: Tian, Zhao, et al.
Published: (2025) -
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
by: Dong, Shijia, et al.
Published: (2026) -
"Refactoring Runaway": Understanding and Mitigating Tangled Refactorings in Coding Agents for Issue Resolution
by: Tian, Zhao, et al.
Published: (2026) -
What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
by: Gu, Wenchao, et al.
Published: (2025)