Holistic Evaluation of State-of-the-Art LLMs for Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Le, Kothari, Suresh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
by: Lang, Nguyet-Anh H., et al.
Published: (2026)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
by: Islam, Md Arman, et al.
Published: (2025)
by: Islam, Md Arman, et al.
Published: (2025)
CodeTF: One-stop Transformer Library for State-of-the-art Code LLMs
by: Bui, Nghi D. Q., et al.
Published: (2023)
by: Bui, Nghi D. Q., et al.
Published: (2023)
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
by: Li, Zenan, et al.
Published: (2026)
by: Li, Zenan, et al.
Published: (2026)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
by: Liang, Shanchao, et al.
Published: (2025)
by: Liang, Shanchao, et al.
Published: (2025)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
by: Pan, Ruwei, et al.
Published: (2026)
by: Pan, Ruwei, et al.
Published: (2026)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
by: Singha, Ananya, et al.
Published: (2025)
by: Singha, Ananya, et al.
Published: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
by: Wang, Hanbin, et al.
Published: (2025)
by: Wang, Hanbin, et al.
Published: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
by: Wang, Shufan, et al.
Published: (2025)
by: Wang, Shufan, et al.
Published: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
by: Li, Ziyu, et al.
Published: (2024)
by: Li, Ziyu, et al.
Published: (2024)
Evaluating SAP Joule for Code Generation
by: Heisler, Joshua, et al.
Published: (2025)
by: Heisler, Joshua, et al.
Published: (2025)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
by: Ho, Anh, et al.
Published: (2025)
by: Ho, Anh, et al.
Published: (2025)
ReCatcher: Towards LLMs Regression Testing for Code Generation
by: Abbassi, Altaf Allah, et al.
Published: (2025)
by: Abbassi, Altaf Allah, et al.
Published: (2025)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
by: Lee, Yunseo, et al.
Published: (2025)
by: Lee, Yunseo, et al.
Published: (2025)
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
by: Sakharova, Marina, et al.
Published: (2025)
by: Sakharova, Marina, et al.
Published: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
by: Wu, Fan, et al.
Published: (2026)
by: Wu, Fan, et al.
Published: (2026)
Function-to-Style Guidance of LLMs for Code Translation
by: Zhang, Longhui, et al.
Published: (2025)
by: Zhang, Longhui, et al.
Published: (2025)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
by: He, Kaifeng, et al.
Published: (2026)
by: He, Kaifeng, et al.
Published: (2026)
Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
by: He, Mengliang, et al.
Published: (2025)
by: He, Mengliang, et al.
Published: (2025)
User Centric Evaluation of Code Generation Tools
by: Miah, Tanha, et al.
Published: (2024)
by: Miah, Tanha, et al.
Published: (2024)
Leveraging Metamemory Mechanisms for Enhanced Data-Free Code Generation in LLMs
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
by: Ngassom, Sylvain Kouemo, et al.
Published: (2024)
A Survey Study on the State of the Art of Programming Exercise Generation using Large Language Models
by: Frankford, Eduard, et al.
Published: (2024)
by: Frankford, Eduard, et al.
Published: (2024)
SpecRover: Code Intent Extraction via LLMs
by: Ruan, Haifeng, et al.
Published: (2024)
by: Ruan, Haifeng, et al.
Published: (2024)
On the Impacts of Contexts on Repository-Level Code Generation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
LiCoEval: Evaluating LLMs on License Compliance in Code Generation
by: Xu, Weiwei, et al.
Published: (2024)
by: Xu, Weiwei, et al.
Published: (2024)
Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
by: Xie, Chen, et al.
Published: (2025)
by: Xie, Chen, et al.
Published: (2025)
Do LLMs Favor Their Providers? Measuring Vertical Integration Bias in Code Generation
by: Catal, Melih, et al.
Published: (2026)
by: Catal, Melih, et al.
Published: (2026)
Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study
by: Chand, Sivajeet, et al.
Published: (2026)
by: Chand, Sivajeet, et al.
Published: (2026)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
by: Deshpande, Ajinkya, et al.
Published: (2024)
by: Deshpande, Ajinkya, et al.
Published: (2024)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
Migrating Code At Scale With LLMs At Google
by: Ziftci, Celal, et al.
Published: (2025)
by: Ziftci, Celal, et al.
Published: (2025)
Analysis on LLMs Performance for Code Summarization
by: Akib, Md. Ahnaf, et al.
Published: (2024)
by: Akib, Md. Ahnaf, et al.
Published: (2024)
Helping LLMs Improve Code Generation Using Feedback from Testing and Static Analysis
by: Dolcetti, Greta, et al.
Published: (2024)
by: Dolcetti, Greta, et al.
Published: (2024)
$\mathbb{USCD}$: Improving Code Generation of LLMs by Uncertainty-Aware Selective Contrastive Decoding
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Similar Items
-
Perish or Flourish? A Holistic Evaluation of Large Language Models for Code Generation in Functional Programming
by: Lang, Nguyet-Anh H., et al.
Published: (2026) -
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
by: Zhang, Lei, et al.
Published: (2026) -
Evaluating the Energy-Efficiency of the Code Generated by LLMs
by: Islam, Md Arman, et al.
Published: (2025) -
CodeTF: One-stop Transformer Library for State-of-the-art Code LLMs
by: Bui, Nghi D. Q., et al.
Published: (2023) -
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
by: Li, Zenan, et al.
Published: (2026)