An evaluation of LLM code generation capabilities through graded exercises
Fuente:
arXiv
Saved in:
| Main Author: | Jiménez, Álvaro Barbero |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025)
by: Singh, Mukul, et al.
Published: (2025)
Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
by: Arefin, Sayed Erfan, et al.
Published: (2023)
by: Arefin, Sayed Erfan, et al.
Published: (2023)
Comparing large language models and human programmers for generating programming code
by: Hou, Wenpin, et al.
Published: (2024)
by: Hou, Wenpin, et al.
Published: (2024)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Assessing LLM code generation quality through path planning tasks
by: Chen, Wanyi, et al.
Published: (2025)
by: Chen, Wanyi, et al.
Published: (2025)
Verification Limits Code LLM Training
by: Gureja, Srishti, et al.
Published: (2025)
by: Gureja, Srishti, et al.
Published: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
by: Tao, Tianhua, et al.
Published: (2024)
by: Tao, Tianhua, et al.
Published: (2024)
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)
by: Cao, Zhuchen, et al.
Published: (2025)
Revolutionizing API Documentation through Summarization
by: Naghshzan, AmirHossein, et al.
Published: (2024)
by: Naghshzan, AmirHossein, et al.
Published: (2024)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
by: Amayuelas, Alfonso, et al.
Published: (2026)
by: Amayuelas, Alfonso, et al.
Published: (2026)
LocAgent: Graph-Guided LLM Agents for Code Localization
by: Chen, Zhaoling, et al.
Published: (2025)
by: Chen, Zhaoling, et al.
Published: (2025)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
by: Fan, Shengda, et al.
Published: (2024)
by: Fan, Shengda, et al.
Published: (2024)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
by: De Michele, William, et al.
Published: (2025)
by: De Michele, William, et al.
Published: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Collaboration is all you need: LLM Assisted Safe Code Translation
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
Issue Localization via LLM-Driven Iterative Code Graph Searching
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
CursorCore: Assist Programming through Aligning Anything
by: Jiang, Hao, et al.
Published: (2024)
by: Jiang, Hao, et al.
Published: (2024)
RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline
by: Zeng, Guancheng, et al.
Published: (2024)
by: Zeng, Guancheng, et al.
Published: (2024)
Using Grammar Masking to Ensure Syntactic Validity in LLM-based Modeling Tasks
by: Netz, Lukas, et al.
Published: (2024)
by: Netz, Lukas, et al.
Published: (2024)
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
by: Xia, Haizhou
Published: (2026)
by: Xia, Haizhou
Published: (2026)
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
by: Shen, Leming, et al.
Published: (2025)
by: Shen, Leming, et al.
Published: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
by: Arnaudo, Anna, et al.
Published: (2026)
by: Arnaudo, Anna, et al.
Published: (2026)
Stingy Context: 18:1 Hierarchical Code Compression for LLM Auto-Coding
by: Ostby, David Linus
Published: (2026)
by: Ostby, David Linus
Published: (2026)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback
by: Hajizadeh, Samira, et al.
Published: (2026)
by: Hajizadeh, Samira, et al.
Published: (2026)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
by: Gupta, Tanmay, et al.
Published: (2024)
by: Gupta, Tanmay, et al.
Published: (2024)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
by: Lee, Seongmin, et al.
Published: (2025)
by: Lee, Seongmin, et al.
Published: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
SoAy: A Solution-based LLM API-using Methodology for Academic Information Seeking
by: Wang, Yuanchun, et al.
Published: (2024)
by: Wang, Yuanchun, et al.
Published: (2024)
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
by: Martinez, Matias, et al.
Published: (2025)
by: Martinez, Matias, et al.
Published: (2025)
Similar Items
-
Diffusion is a code repair operator and generator
by: Singh, Mukul, et al.
Published: (2025) -
Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
by: Arefin, Sayed Erfan, et al.
Published: (2023) -
Comparing large language models and human programmers for generating programming code
by: Hou, Wenpin, et al.
Published: (2024) -
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
by: Zhang, Tianyi, et al.
Published: (2025) -
Assessing LLM code generation quality through path planning tasks
by: Chen, Wanyi, et al.
Published: (2025)