Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | Chi, Wayne, Chen, Valerie, Angelopoulos, Anastasios Nikolas, Chiang, Wei-Lin, Mittal, Aditya, Jain, Naman, Zhang, Tianjun, Stoica, Ion, Donahue, Chris, Talwalkar, Ameet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
Comparing Developer and LLM Biases in Code Evaluation
by: Mittal, Aditya, et al.
Published: (2026)
by: Mittal, Aditya, et al.
Published: (2026)
The Impact of Element Ordering on LM Agent Performance
by: Chi, Wayne, et al.
Published: (2024)
by: Chi, Wayne, et al.
Published: (2024)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
by: Chen, Valerie, et al.
Published: (2025)
by: Chen, Valerie, et al.
Published: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)
by: Chen, Valerie, et al.
Published: (2026)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
by: Jain, Naman, et al.
Published: (2024)
by: Jain, Naman, et al.
Published: (2024)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation
by: Jiang, Mingchao, et al.
Published: (2025)
by: Jiang, Mingchao, et al.
Published: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
by: Shetty, Manish, et al.
Published: (2025)
by: Shetty, Manish, et al.
Published: (2025)
Music Arena: Live Evaluation for Text-to-Music
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
by: Chiang, Wei-Lin, et al.
Published: (2024)
by: Chiang, Wei-Lin, et al.
Published: (2024)
Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using Copilot
by: Zhang, Beiqi, et al.
Published: (2024)
by: Zhang, Beiqi, et al.
Published: (2024)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
by: He, Keyu, et al.
Published: (2026)
by: He, Keyu, et al.
Published: (2026)
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
GitHub Copilot: the perfect Code compLeeter?
by: Siroš, Ilja, et al.
Published: (2024)
by: Siroš, Ilja, et al.
Published: (2024)
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
by: Peng, Wenjun, et al.
Published: (2026)
by: Peng, Wenjun, et al.
Published: (2026)
Synthesizing Performance Constraints for Evaluating and Improving Code Efficiency
by: Yang, Jun, et al.
Published: (2025)
by: Yang, Jun, et al.
Published: (2025)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
by: Mozannar, Hussein, et al.
Published: (2024)
by: Mozannar, Hussein, et al.
Published: (2024)
Is GitHub's Copilot as Bad as Humans at Introducing Vulnerabilities in Code?
by: Asare, Owura, et al.
Published: (2022)
by: Asare, Owura, et al.
Published: (2022)
Measuring the Runtime Performance of C++ Code Written by Humans using GitHub Copilot
by: Erhabor, Daniel, et al.
Published: (2023)
by: Erhabor, Daniel, et al.
Published: (2023)
Exploring the Effect of Multiple Natural Languages on Code Suggestion Using GitHub Copilot
by: Koyanagi, Kei, et al.
Published: (2024)
by: Koyanagi, Kei, et al.
Published: (2024)
Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
by: Zhang, Binquan, et al.
Published: (2025)
by: Zhang, Binquan, et al.
Published: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
by: Pathak, Aditya, et al.
Published: (2025)
by: Pathak, Aditya, et al.
Published: (2025)
Figma2Code: Automating Multimodal Design to Code in the Wild
by: Gui, Yi, et al.
Published: (2026)
by: Gui, Yi, et al.
Published: (2026)
Articulate but Wrong: Self-Review Failures in LLM-Based Code Modernization
by: Reddy, Gokul Chandra Purnachandra, et al.
Published: (2026)
by: Reddy, Gokul Chandra Purnachandra, et al.
Published: (2026)
Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical Study
by: Fu, Yujia, et al.
Published: (2023)
by: Fu, Yujia, et al.
Published: (2023)
CodingGenie: A Proactive LLM-Powered Programming Assistant
by: Zhao, Sebastian, et al.
Published: (2025)
by: Zhao, Sebastian, et al.
Published: (2025)
Code Comprehension with GitHub Copilot: Performance Gains, Comprehension Trade-offs, and Behavioral Predictors in Brownfield Programming
by: Qiao, Yunhan, et al.
Published: (2025)
by: Qiao, Yunhan, et al.
Published: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
by: Zhao, Zhimin
Published: (2025)
by: Zhao, Zhimin
Published: (2025)
Programming with AI: Evaluating ChatGPT, Gemini, AlphaCode, and GitHub Copilot for Programmers
by: Siam, Md Kamrul, et al.
Published: (2024)
by: Siam, Md Kamrul, et al.
Published: (2024)
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
by: Borovkov, Nikita, et al.
Published: (2026)
by: Borovkov, Nikita, et al.
Published: (2026)
From Code Generation to Software Testing: AI Copilot with Context-Based RAG
by: Wang, Yuchen, et al.
Published: (2025)
by: Wang, Yuchen, et al.
Published: (2025)
RevMine: An LLM-Assisted Tool for Code Review Mining and Analysis Across Git Platforms
by: Kansab, Samah, et al.
Published: (2025)
by: Kansab, Samah, et al.
Published: (2025)
WildCode: An Empirical Analysis of Code Generated by ChatGPT
by: Khanmohammadi, Kobra, et al.
Published: (2025)
by: Khanmohammadi, Kobra, et al.
Published: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
by: Liu, Zeyu Leo, et al.
Published: (2024)
by: Liu, Zeyu Leo, et al.
Published: (2024)
ECO: An LLM-Driven Efficient Code Optimizer for Warehouse Scale Computers
by: Lin, Hannah, et al.
Published: (2025)
by: Lin, Hannah, et al.
Published: (2025)
Similar Items
-
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025) -
Comparing Developer and LLM Biases in Code Evaluation
by: Mittal, Aditya, et al.
Published: (2026) -
The Impact of Element Ordering on LM Agent Performance
by: Chi, Wayne, et al.
Published: (2024) -
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
by: Chen, Valerie, et al.
Published: (2025) -
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
by: Chen, Valerie, et al.
Published: (2026)