Wisdom and Delusion of LLM Ensembles for Code Generation and Repair
Fuente:
arXiv
Guardado en:
| Autores principales: | Vallecillos-Ruiz, Fernando, Hort, Max, Moonen, Leon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2025)
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2025)
Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2024)
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2024)
Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
por: Hort, Max, et al.
Publicado: (2025)
por: Hort, Max, et al.
Publicado: (2025)
The Impact of Fine-tuning Large Language Models on Automated Program Repair
por: Macháček, Roman, et al.
Publicado: (2025)
por: Macháček, Roman, et al.
Publicado: (2025)
Semantic-Preserving Transformations as Mutation Operators: A Study on Their Effectiveness in Defect Detection
por: Hort, Max, et al.
Publicado: (2025)
por: Hort, Max, et al.
Publicado: (2025)
A Comparative Study on Large Language Models for Log Parsing
por: Astekin, Merve, et al.
Publicado: (2024)
por: Astekin, Merve, et al.
Publicado: (2024)
A Semantic-based Optimization Approach for Repairing LLMs: Case Study on Code Generation
por: Gu, Jian, et al.
Publicado: (2025)
por: Gu, Jian, et al.
Publicado: (2025)
Neuron Patching: Semantic-based Neuron-level Language Model Repair for Code Generation
por: Gu, Jian, et al.
Publicado: (2023)
por: Gu, Jian, et al.
Publicado: (2023)
Aligning the Objective of LLM-based Program Repair
por: Xu, Junjielong, et al.
Publicado: (2024)
por: Xu, Junjielong, et al.
Publicado: (2024)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
por: Peng, Jinjun, et al.
Publicado: (2025)
por: Peng, Jinjun, et al.
Publicado: (2025)
IntentCoding: Amplifying User Intent in Code Generation
por: Fang, Zheng, et al.
Publicado: (2026)
por: Fang, Zheng, et al.
Publicado: (2026)
SelfCodeAlign: Self-Alignment for Code Generation
por: Wei, Yuxiang, et al.
Publicado: (2024)
por: Wei, Yuxiang, et al.
Publicado: (2024)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
por: Weyssow, Martin, et al.
Publicado: (2024)
por: Weyssow, Martin, et al.
Publicado: (2024)
CodeJudge: Evaluating Code Generation with Large Language Models
por: Tong, Weixi, et al.
Publicado: (2024)
por: Tong, Weixi, et al.
Publicado: (2024)
AuPair: Golden Example Pairs for Code Repair
por: Mavalankar, Aditi, et al.
Publicado: (2025)
por: Mavalankar, Aditi, et al.
Publicado: (2025)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
por: Huang, Jue, et al.
Publicado: (2026)
por: Huang, Jue, et al.
Publicado: (2026)
Evaluating Language Models for Efficient Code Generation
por: Liu, Jiawei, et al.
Publicado: (2024)
por: Liu, Jiawei, et al.
Publicado: (2024)
A Survey on Code Generation with LLM-based Agents
por: Dong, Yihong, et al.
Publicado: (2025)
por: Dong, Yihong, et al.
Publicado: (2025)
GiFT: Gibbs Fine-Tuning for Code Generation
por: Li, Haochen, et al.
Publicado: (2025)
por: Li, Haochen, et al.
Publicado: (2025)
MetaLint: Easy-to-Hard Generalization for Code Linting
por: Naik, Atharva, et al.
Publicado: (2025)
por: Naik, Atharva, et al.
Publicado: (2025)
GrowthHacker: Automated Off-Policy Evaluation Optimization Using Code-Modifying LLM Agents
por: Wu, Jie JW, et al.
Publicado: (2025)
por: Wu, Jie JW, et al.
Publicado: (2025)
Test Code Generation for Telecom Software Systems using Two-Stage Generative Model
por: Nabeel, Mohamad, et al.
Publicado: (2024)
por: Nabeel, Mohamad, et al.
Publicado: (2024)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
por: Yang, Rem, et al.
Publicado: (2025)
por: Yang, Rem, et al.
Publicado: (2025)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
por: Xie, Yiqing, et al.
Publicado: (2026)
por: Xie, Yiqing, et al.
Publicado: (2026)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
por: Gong, Linyuan, et al.
Publicado: (2024)
por: Gong, Linyuan, et al.
Publicado: (2024)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
por: Riddell, Martin, et al.
Publicado: (2024)
por: Riddell, Martin, et al.
Publicado: (2024)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
por: Ridnik, Tal, et al.
Publicado: (2024)
por: Ridnik, Tal, et al.
Publicado: (2024)
SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
por: Xia, Xiao, et al.
Publicado: (2024)
por: Xia, Xiao, et al.
Publicado: (2024)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
por: Hassid, Michael, et al.
Publicado: (2024)
por: Hassid, Michael, et al.
Publicado: (2024)
Agent-Driven Automatic Software Improvement
por: Ruiz, Fernando Vallecillos
Publicado: (2024)
por: Ruiz, Fernando Vallecillos
Publicado: (2024)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
por: Haider, Md. Asif, et al.
Publicado: (2024)
por: Haider, Md. Asif, et al.
Publicado: (2024)
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
por: Weyssow, Martin, et al.
Publicado: (2023)
por: Weyssow, Martin, et al.
Publicado: (2023)
CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing
por: He, Xinyi, et al.
Publicado: (2024)
por: He, Xinyi, et al.
Publicado: (2024)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
por: Petrukha, Ivan, et al.
Publicado: (2025)
por: Petrukha, Ivan, et al.
Publicado: (2025)
GLLM: Self-Corrective G-Code Generation using Large Language Models with User Feedback
por: Abdelaal, Mohamed, et al.
Publicado: (2025)
por: Abdelaal, Mohamed, et al.
Publicado: (2025)
Let the Code LLM Edit Itself When You Edit the Code
por: He, Zhenyu, et al.
Publicado: (2024)
por: He, Zhenyu, et al.
Publicado: (2024)
TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Models
por: Park, Chansung, et al.
Publicado: (2026)
por: Park, Chansung, et al.
Publicado: (2026)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
por: Jain, Naman, et al.
Publicado: (2024)
por: Jain, Naman, et al.
Publicado: (2024)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
por: Jiang, Juyong, et al.
Publicado: (2026)
por: Jiang, Juyong, et al.
Publicado: (2026)
Ejemplares similares
-
The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2025) -
Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
por: Ruiz, Fernando Vallecillos, et al.
Publicado: (2024) -
Codehacks: A Dataset of Adversarial Tests for Competitive Programming Problems Obtained from Codeforces
por: Hort, Max, et al.
Publicado: (2025) -
The Impact of Fine-tuning Large Language Models on Automated Program Repair
por: Macháček, Roman, et al.
Publicado: (2025) -
Semantic-Preserving Transformations as Mutation Operators: A Study on Their Effectiveness in Defect Detection
por: Hort, Max, et al.
Publicado: (2025)