When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pan, Jane, Shar, Ryan, Pfau, Jacob, Talwalkar, Ameet, He, He, Chen, Valerie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915172628561920
author Pan, Jane
Shar, Ryan
Pfau, Jacob
Talwalkar, Ameet
He, He
Chen, Valerie
author_facet Pan, Jane
Shar, Ryan
Pfau, Jacob
Talwalkar, Ameet
He, He
Chen, Valerie
contents Programming is a fundamentally interactive process, yet coding assistants are often evaluated using static benchmarks that fail to measure how well models collaborate with users. We introduce an interactive evaluation pipeline to examine how LLMs incorporate different types of feedback in a collaborative setting. Specifically, we perturb static coding benchmarks so that the code model must interact with a simulated user to retrieve key information about the problem. We find that interaction significantly affects model performance, as the relative rankings of 10 models across 3 datasets often vary between static and interactive settings, despite models being fairly robust to feedback that contains errors. We also observe that even when different feedback types are equally effective with respect to performance, they can impact model behaviors such as (1) how models respond to higher- vs. lower-quality feedback and (2) whether models prioritize aesthetic vs. functional edits. Our work aims to "re-evaluate" model coding capabilities through an interactive lens toward bridging the gap between existing evaluations and real-world usage.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18413
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
Pan, Jane
Shar, Ryan
Pfau, Jacob
Talwalkar, Ameet
He, He
Chen, Valerie
Human-Computer Interaction
Programming is a fundamentally interactive process, yet coding assistants are often evaluated using static benchmarks that fail to measure how well models collaborate with users. We introduce an interactive evaluation pipeline to examine how LLMs incorporate different types of feedback in a collaborative setting. Specifically, we perturb static coding benchmarks so that the code model must interact with a simulated user to retrieve key information about the problem. We find that interaction significantly affects model performance, as the relative rankings of 10 models across 3 datasets often vary between static and interactive settings, despite models being fairly robust to feedback that contains errors. We also observe that even when different feedback types are equally effective with respect to performance, they can impact model behaviors such as (1) how models respond to higher- vs. lower-quality feedback and (2) whether models prioritize aesthetic vs. functional edits. Our work aims to "re-evaluate" model coding capabilities through an interactive lens toward bridging the gap between existing evaluations and real-world usage.
title When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
topic Human-Computer Interaction
url https://arxiv.org/abs/2502.18413