CodeArena: A Collective Evaluation Platform for LLM Code Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Du, Mingzhe, Luu, Anh Tuan, Ji, Bin, Wu, Xiaobao, Huang, Dong, Zhuo, Terry Yue, Liu, Qian, Ng, See-Kiong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mercury: A Code Efficiency Benchmark for Code Large Language Models
di: Du, Mingzhe, et al.
Pubblicazione: (2024)
di: Du, Mingzhe, et al.
Pubblicazione: (2024)
Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
di: Du, Mingzhe, et al.
Pubblicazione: (2025)
di: Du, Mingzhe, et al.
Pubblicazione: (2025)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
di: Chi, Wayne, et al.
Pubblicazione: (2025)
di: Chi, Wayne, et al.
Pubblicazione: (2025)
ICE-Score: Instructing Large Language Models to Evaluate Code
di: Zhuo, Terry Yue
Pubblicazione: (2023)
di: Zhuo, Terry Yue
Pubblicazione: (2023)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
CodeScore: Evaluating Code Generation by Learning Code Execution
di: Dong, Yihong, et al.
Pubblicazione: (2023)
di: Dong, Yihong, et al.
Pubblicazione: (2023)
From Code to Courtroom: LLMs as the New Software Judges
di: He, Junda, et al.
Pubblicazione: (2025)
di: He, Junda, et al.
Pubblicazione: (2025)
On Evaluating the Efficiency of Source Code Generated by LLMs
di: Niu, Changan, et al.
Pubblicazione: (2024)
di: Niu, Changan, et al.
Pubblicazione: (2024)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
di: Yang, Zhou, et al.
Pubblicazione: (2024)
di: Yang, Zhou, et al.
Pubblicazione: (2024)
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models
di: Yang, Guang, et al.
Pubblicazione: (2023)
di: Yang, Guang, et al.
Pubblicazione: (2023)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
Measuring the Influence of Incorrect Code on Test Generation
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
CodeLSI: Leveraging Foundation Models for Automated Code Generation with Low-Rank Optimization and Domain-Specific Instruction Tuning
di: Le, Huy, et al.
Pubblicazione: (2025)
di: Le, Huy, et al.
Pubblicazione: (2025)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
di: He, Xinyi, et al.
Pubblicazione: (2025)
di: He, Xinyi, et al.
Pubblicazione: (2025)
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
di: Peng, Wenjun, et al.
Pubblicazione: (2026)
di: Peng, Wenjun, et al.
Pubblicazione: (2026)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
di: Gong, Zhihao, et al.
Pubblicazione: (2026)
di: Gong, Zhihao, et al.
Pubblicazione: (2026)
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation
di: Gong, Zhihao, et al.
Pubblicazione: (2025)
di: Gong, Zhihao, et al.
Pubblicazione: (2025)
Less is More: DocString Compression in Code Generation
di: Yang, Guang, et al.
Pubblicazione: (2024)
di: Yang, Guang, et al.
Pubblicazione: (2024)
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2025)
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2025)
Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code Rewriting
di: Ye, Tong, et al.
Pubblicazione: (2024)
di: Ye, Tong, et al.
Pubblicazione: (2024)
Coherence Collapse: Diagnosing Why Code Agents Fail After Reaching the Right Code
di: Kim, Myeongsoo, et al.
Pubblicazione: (2026)
di: Kim, Myeongsoo, et al.
Pubblicazione: (2026)
CodeWiki: Evaluating AI's Ability to Generate Holistic Documentation for Large-Scale Codebases
di: Hoang, Anh Nguyen, et al.
Pubblicazione: (2025)
di: Hoang, Anh Nguyen, et al.
Pubblicazione: (2025)
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
di: Dong, Shijia, et al.
Pubblicazione: (2026)
di: Dong, Shijia, et al.
Pubblicazione: (2026)
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
di: Zhang, Jiajun, et al.
Pubblicazione: (2026)
di: Zhang, Jiajun, et al.
Pubblicazione: (2026)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
di: Islah, Nizar, et al.
Pubblicazione: (2024)
di: Islah, Nizar, et al.
Pubblicazione: (2024)
Inducing Vulnerable Code Generation in LLM Coding Assistants
di: Zeng, Binqi, et al.
Pubblicazione: (2025)
di: Zeng, Binqi, et al.
Pubblicazione: (2025)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
di: Dau, Anh T. V., et al.
Pubblicazione: (2026)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
di: Yang, Guang, et al.
Pubblicazione: (2024)
di: Yang, Guang, et al.
Pubblicazione: (2024)
Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2024)
di: Zhuo, Terry Yue, et al.
Pubblicazione: (2024)
RevMine: An LLM-Assisted Tool for Code Review Mining and Analysis Across Git Platforms
di: Kansab, Samah, et al.
Pubblicazione: (2025)
di: Kansab, Samah, et al.
Pubblicazione: (2025)
Evaluating LLM-Generated Code: A Benchmark and Developer Study
di: Szych, Joanna, et al.
Pubblicazione: (2026)
di: Szych, Joanna, et al.
Pubblicazione: (2026)
Evaluating Efficiency and Novelty of LLM-Generated Code for Graph Analysis
di: Nia, Atieh Barati, et al.
Pubblicazione: (2025)
di: Nia, Atieh Barati, et al.
Pubblicazione: (2025)
UniCode: Augmenting Evaluation for Code Reasoning
di: Zheng, Xinyue, et al.
Pubblicazione: (2025)
di: Zheng, Xinyue, et al.
Pubblicazione: (2025)
Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
di: Vu, Thanh Trong, et al.
Pubblicazione: (2025)
di: Vu, Thanh Trong, et al.
Pubblicazione: (2025)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
di: He, Fusen, et al.
Pubblicazione: (2024)
di: He, Fusen, et al.
Pubblicazione: (2024)
Context-Aware CodeLLM Eviction for AI-assisted Coding
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
di: Thangarajah, Kishanthan, et al.
Pubblicazione: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
di: He, Junda, et al.
Pubblicazione: (2025)
di: He, Junda, et al.
Pubblicazione: (2025)
DocChecker: Bootstrapping Code Large Language Model for Detecting and Resolving Code-Comment Inconsistencies
di: Dau, Anh T. V., et al.
Pubblicazione: (2023)
di: Dau, Anh T. V., et al.
Pubblicazione: (2023)
Documenti analoghi
-
Mercury: A Code Efficiency Benchmark for Code Large Language Models
di: Du, Mingzhe, et al.
Pubblicazione: (2024) -
Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
di: Du, Mingzhe, et al.
Pubblicazione: (2025) -
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
di: Chi, Wayne, et al.
Pubblicazione: (2025) -
ICE-Score: Instructing Large Language Models to Evaluate Code
di: Zhuo, Terry Yue
Pubblicazione: (2023) -
Benchmarking LLMs for Unit Test Generation from Real-World Functions
di: Huang, Dong, et al.
Pubblicazione: (2025)