EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chi, Wayne, Chen, Valerie, Shar, Ryan, Mittal, Aditya, Liang, Jenny, Chiang, Wei-Lin, Angelopoulos, Anastasios Nikolas, Stoica, Ion, Neubig, Graham, Talwalkar, Ameet, Donahue, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
von: Chi, Wayne, et al.
Veröffentlicht: (2025)
The Impact of Element Ordering on LM Agent Performance
von: Chi, Wayne, et al.
Veröffentlicht: (2024)
von: Chi, Wayne, et al.
Veröffentlicht: (2024)
Comparing Developer and LLM Biases in Code Evaluation
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
von: Mittal, Aditya, et al.
Veröffentlicht: (2026)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
Do LLMs exhibit human-like response biases? A case study in survey design
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
von: Pan, Jane, et al.
Veröffentlicht: (2025)
von: Pan, Jane, et al.
Veröffentlicht: (2025)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
Music Arena: Live Evaluation for Text-to-Music
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
von: Kim, Yonghyun, et al.
Veröffentlicht: (2025)
Multitask Learning Can Improve Worst-Group Outcomes
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2023)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2023)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
von: Yu, Rebecca, et al.
Veröffentlicht: (2025)
von: Yu, Rebecca, et al.
Veröffentlicht: (2025)
Prompt-to-Leaderboard
von: Frick, Evan, et al.
Veröffentlicht: (2025)
von: Frick, Evan, et al.
Veröffentlicht: (2025)
How can we assess human-agent interactions? Case studies in software agent design
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
von: Chen, Valerie, et al.
Veröffentlicht: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
von: Shen, Junhong, et al.
Veröffentlicht: (2024)
von: Shen, Junhong, et al.
Veröffentlicht: (2024)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
von: Chen, Valerie, et al.
Veröffentlicht: (2026)
von: Chen, Valerie, et al.
Veröffentlicht: (2026)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization
von: Mishra, Prakamya, et al.
Veröffentlicht: (2024)
von: Mishra, Prakamya, et al.
Veröffentlicht: (2024)
Some Present-Day Problems of Romanian Library Science
von: Stoica, Ion
Veröffentlicht: (1973)
von: Stoica, Ion
Veröffentlicht: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
von: Stoica, Ion
Veröffentlicht: (1972)
von: Stoica, Ion
Veröffentlicht: (1972)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
von: Mozannar, Hussein, et al.
Veröffentlicht: (2024)
von: Mozannar, Hussein, et al.
Veröffentlicht: (2024)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Need Help? Designing Proactive AI Assistants for Programming
von: Chen, Valerie, et al.
Veröffentlicht: (2024)
von: Chen, Valerie, et al.
Veröffentlicht: (2024)
CodingGenie: A Proactive LLM-Powered Programming Assistant
von: Zhao, Sebastian, et al.
Veröffentlicht: (2025)
von: Zhao, Sebastian, et al.
Veröffentlicht: (2025)
Conformal Risk Control for Non-Monotonic Losses
von: Angelopoulos, Anastasios N.
Veröffentlicht: (2026)
von: Angelopoulos, Anastasios N.
Veröffentlicht: (2026)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
von: Huang, Baihe, et al.
Veröffentlicht: (2025)
von: Huang, Baihe, et al.
Veröffentlicht: (2025)
Agreement-Based Cascading for Efficient Inference
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
von: Kolawole, Steven, et al.
Veröffentlicht: (2024)
EDIT0RIAL
von: La Dirección
Veröffentlicht: (2007)
von: La Dirección
Veröffentlicht: (2007)
Provably tuning the ElasticNet across instances
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2022)
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2022)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
von: Khodak, Mikhail, et al.
Veröffentlicht: (2023)
von: Khodak, Mikhail, et al.
Veröffentlicht: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
von: Johnson, Nari, et al.
Veröffentlicht: (2023)
von: Johnson, Nari, et al.
Veröffentlicht: (2023)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
von: Xie, Stephan, et al.
Veröffentlicht: (2026)
von: Xie, Stephan, et al.
Veröffentlicht: (2026)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
von: Soni, Aditya Bharat, et al.
Veröffentlicht: (2025)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
von: Li, Sijie, et al.
Veröffentlicht: (2025)
von: Li, Sijie, et al.
Veröffentlicht: (2025)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
von: Feng, Shengyu, et al.
Veröffentlicht: (2025)
von: Feng, Shengyu, et al.
Veröffentlicht: (2025)
Specifications: The missing link to making the development of LLM systems an engineering discipline
von: Stoica, Ion, et al.
Veröffentlicht: (2024)
von: Stoica, Ion, et al.
Veröffentlicht: (2024)
Learning Personalized Decision Support Policies
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
von: Bhatt, Umang, et al.
Veröffentlicht: (2023)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
von: Li, Tianle, et al.
Veröffentlicht: (2024)
von: Li, Tianle, et al.
Veröffentlicht: (2024)
CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
von: Mehralian, Forough, et al.
Veröffentlicht: (2025)
von: Mehralian, Forough, et al.
Veröffentlicht: (2025)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
von: Wang, Xingyao, et al.
Veröffentlicht: (2026)
von: Wang, Xingyao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
von: Chi, Wayne, et al.
Veröffentlicht: (2025) -
The Impact of Element Ordering on LM Agent Performance
von: Chi, Wayne, et al.
Veröffentlicht: (2024) -
Comparing Developer and LLM Biases in Code Evaluation
von: Mittal, Aditya, et al.
Veröffentlicht: (2026) -
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
von: Chen, Valerie, et al.
Veröffentlicht: (2025) -
Do LLMs exhibit human-like response biases? A case study in survey design
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2023)