Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
Fuente:
arXiv
Salvato in:
| Autori principali: | Tsai, Yun-Da, Liu, Mingjie, Ren, Haoxing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
di: Gao, Pengfei, et al.
Pubblicazione: (2025)
di: Gao, Pengfei, et al.
Pubblicazione: (2025)
SelfCodeAlign: Self-Alignment for Code Generation
di: Wei, Yuxiang, et al.
Pubblicazione: (2024)
di: Wei, Yuxiang, et al.
Pubblicazione: (2024)
Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
di: Wang, Changsheng, et al.
Pubblicazione: (2025)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
di: Haider, Md. Asif, et al.
Pubblicazione: (2024)
di: Haider, Md. Asif, et al.
Pubblicazione: (2024)
Code Roulette: How Prompt Variability Affects LLM Code Generation
di: Paleyes, Andrei, et al.
Pubblicazione: (2025)
di: Paleyes, Andrei, et al.
Pubblicazione: (2025)
UniASM: Binary Code Similarity Detection without Fine-tuning
di: Gu, Yeming, et al.
Pubblicazione: (2022)
di: Gu, Yeming, et al.
Pubblicazione: (2022)
LLM Performance for Code Generation on Noisy Tasks
di: Sendyka, Radzim, et al.
Pubblicazione: (2025)
di: Sendyka, Radzim, et al.
Pubblicazione: (2025)
Renaissance of Literate Programming in the Era of LLMs: Enhancing LLM-Based Code Generation in Large-Scale Projects
di: Zhang, Wuyang, et al.
Pubblicazione: (2024)
di: Zhang, Wuyang, et al.
Pubblicazione: (2024)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
di: Weyssow, Martin, et al.
Pubblicazione: (2024)
di: Weyssow, Martin, et al.
Pubblicazione: (2024)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
di: Park, Chansung, et al.
Pubblicazione: (2026)
di: Park, Chansung, et al.
Pubblicazione: (2026)
Less is More: DocString Compression in Code Generation
di: Yang, Guang, et al.
Pubblicazione: (2024)
di: Yang, Guang, et al.
Pubblicazione: (2024)
Refining Joint Text and Source Code Embeddings for Retrieval Task with Parameter-Efficient Fine-Tuning
di: Galliamov, Karim, et al.
Pubblicazione: (2024)
di: Galliamov, Karim, et al.
Pubblicazione: (2024)
JARVIS: A Multi-Agent Code Assistant for High-Quality EDA Script Generation
di: Pasandi, Ghasem, et al.
Pubblicazione: (2025)
di: Pasandi, Ghasem, et al.
Pubblicazione: (2025)
Think Anywhere in Code Generation
di: Jiang, Xue, et al.
Pubblicazione: (2026)
di: Jiang, Xue, et al.
Pubblicazione: (2026)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
di: Li, Xin-Ye, et al.
Pubblicazione: (2026)
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
di: Weyssow, Martin, et al.
Pubblicazione: (2023)
di: Weyssow, Martin, et al.
Pubblicazione: (2023)
SemRep: Generative Code Representation Learning with Code Transformations
di: Li, Weichen, et al.
Pubblicazione: (2026)
di: Li, Weichen, et al.
Pubblicazione: (2026)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
di: Diggs, Colin, et al.
Pubblicazione: (2024)
di: Diggs, Colin, et al.
Pubblicazione: (2024)
GiFT: Gibbs Fine-Tuning for Code Generation
di: Li, Haochen, et al.
Pubblicazione: (2025)
di: Li, Haochen, et al.
Pubblicazione: (2025)
Evaluating Language Models for Efficient Code Generation
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
di: Liu, Jiawei, et al.
Pubblicazione: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
di: Qiu, Ruizhong, et al.
Pubblicazione: (2024)
Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
di: Zeng, Wenhao, et al.
Pubblicazione: (2025)
di: Zeng, Wenhao, et al.
Pubblicazione: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
di: Xu, WeiZhe, et al.
Pubblicazione: (2026)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
di: Yan, Kaiwen, et al.
Pubblicazione: (2025)
di: Yan, Kaiwen, et al.
Pubblicazione: (2025)
Timing Analysis Agent: Autonomous Multi-Corner Multi-Mode (MCMM) Timing Debugging with Timing Debug Relation Graph
di: Nainani, Jatin, et al.
Pubblicazione: (2025)
di: Nainani, Jatin, et al.
Pubblicazione: (2025)
A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
di: Joel, Sathvik, et al.
Pubblicazione: (2024)
di: Joel, Sathvik, et al.
Pubblicazione: (2024)
Leveraging Reviewer Experience in Code Review Comment Generation
di: Lin, Hong Yi, et al.
Pubblicazione: (2024)
di: Lin, Hong Yi, et al.
Pubblicazione: (2024)
When Less is More: On the Value of "Co-training" for Semi-Supervised Software Defect Predictors
di: Majumder, Suvodeep, et al.
Pubblicazione: (2022)
di: Majumder, Suvodeep, et al.
Pubblicazione: (2022)
Wisdom and Delusion of LLM Ensembles for Code Generation and Repair
di: Vallecillos-Ruiz, Fernando, et al.
Pubblicazione: (2025)
di: Vallecillos-Ruiz, Fernando, et al.
Pubblicazione: (2025)
Vibe Checker: Aligning Code Evaluation with Human Preference
di: Zhong, Ming, et al.
Pubblicazione: (2025)
di: Zhong, Ming, et al.
Pubblicazione: (2025)
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
di: Ryan, Gabriel, et al.
Pubblicazione: (2024)
di: Ryan, Gabriel, et al.
Pubblicazione: (2024)
Fault Localization via Fine-tuning Large Language Models with Mutation Generated Stack Traces
di: Jambigi, Neetha, et al.
Pubblicazione: (2025)
di: Jambigi, Neetha, et al.
Pubblicazione: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
di: Akli, Amal, et al.
Pubblicazione: (2026)
di: Akli, Amal, et al.
Pubblicazione: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
di: Jiang, Shan, et al.
Pubblicazione: (2026)
di: Jiang, Shan, et al.
Pubblicazione: (2026)
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View Graphs
di: Mathai, Alex, et al.
Pubblicazione: (2024)
di: Mathai, Alex, et al.
Pubblicazione: (2024)
StructCoder: Structure-Aware Transformer for Code Generation
di: Tipirneni, Sindhu, et al.
Pubblicazione: (2022)
di: Tipirneni, Sindhu, et al.
Pubblicazione: (2022)
Less is More: Towards Green Code Large Language Models via Unified Structural Pruning
di: Yang, Guang, et al.
Pubblicazione: (2024)
di: Yang, Guang, et al.
Pubblicazione: (2024)
IntentCoding: Amplifying User Intent in Code Generation
di: Fang, Zheng, et al.
Pubblicazione: (2026)
di: Fang, Zheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
di: Gao, Pengfei, et al.
Pubblicazione: (2025) -
SelfCodeAlign: Self-Alignment for Code Generation
di: Wei, Yuxiang, et al.
Pubblicazione: (2024) -
Breaking Memorization Barriers in LLM Code Fine-Tuning via Information Bottleneck for Improved Generalization
di: Wang, Changsheng, et al.
Pubblicazione: (2025) -
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
di: Haider, Md. Asif, et al.
Pubblicazione: (2024) -
Code Roulette: How Prompt Variability Affects LLM Code Generation
di: Paleyes, Andrei, et al.
Pubblicazione: (2025)