RubberDuckBench: A Benchmark for AI Coding Assistants

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mohammed, Ferida, Ayad, Fatma, Maniatis, Petros, Chandra, Satish, Dinella, Elizabeth
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909012948156416
author Mohammed, Ferida
Ayad, Fatma
Maniatis, Petros
Chandra, Satish
Dinella, Elizabeth
author_facet Mohammed, Ferida
Ayad, Fatma
Maniatis, Petros
Chandra, Satish
Dinella, Elizabeth
contents Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To enable such a study, we curate a benchmark of real-world contextualized questions derived from Github pull request comments. Out of this work, we present RubberDuckBench: a multilingual benchmark of questions about code, along with detailed rubrics for evaluating answers. We evaluate a diverse set of 20 LLMs (proprietary & open-source) on answering these questions. We find that even state of the art models fail to give consistent, correct responses across the benchmark. Grok 4 (69.29%), Claude Opus 4 (68.5%), and GPT-5 (67.8%) perform best overall, but do not exhibit pairwise significant superiority over the next 9 best performing models. Most models obtain points through partial credit, with the best performing models only answering at most 2 questions completely correctly across all trials. Furthermore, models often hallucinate with lies in 58.3\% of responses on average. Cost analysis reveals no correlation between expense (API pricing or parameter count) and performance. We intend this benchmark to be a target for future research in trustworthy and correct AI coding assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16456
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RubberDuckBench: A Benchmark for AI Coding Assistants
Mohammed, Ferida
Ayad, Fatma
Maniatis, Petros
Chandra, Satish
Dinella, Elizabeth
Software Engineering
Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To enable such a study, we curate a benchmark of real-world contextualized questions derived from Github pull request comments. Out of this work, we present RubberDuckBench: a multilingual benchmark of questions about code, along with detailed rubrics for evaluating answers. We evaluate a diverse set of 20 LLMs (proprietary & open-source) on answering these questions. We find that even state of the art models fail to give consistent, correct responses across the benchmark. Grok 4 (69.29%), Claude Opus 4 (68.5%), and GPT-5 (67.8%) perform best overall, but do not exhibit pairwise significant superiority over the next 9 best performing models. Most models obtain points through partial credit, with the best performing models only answering at most 2 questions completely correctly across all trials. Furthermore, models often hallucinate with lies in 58.3\% of responses on average. Cost analysis reveals no correlation between expense (API pricing or parameter count) and performance. We intend this benchmark to be a target for future research in trustworthy and correct AI coding assistants.
title RubberDuckBench: A Benchmark for AI Coding Assistants
topic Software Engineering
url https://arxiv.org/abs/2601.16456