RubberDuckBench: A Benchmark for AI Coding Assistants
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohammed, Ferida, Ayad, Fatma, Maniatis, Petros, Chandra, Satish, Dinella, Elizabeth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRQBench: A Benchmark of Code Reasoning Questions
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2024)
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2024)
On the Effectiveness of Modular Testing in EvoSuite
von: Dinella, Elizabeth
Veröffentlicht: (2026)
von: Dinella, Elizabeth
Veröffentlicht: (2026)
Program Structure Aware Precondition Generation
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2023)
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2023)
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
von: Mathai, Alex, et al.
Veröffentlicht: (2024)
von: Mathai, Alex, et al.
Veröffentlicht: (2024)
Agentic Code Reasoning
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
von: Huang, Chenxi, et al.
Veröffentlicht: (2026)
von: Huang, Chenxi, et al.
Veröffentlicht: (2026)
AI-Assisted Assessment of Coding Practices in Modern Code Review
von: Vijayvergiya, Manushree, et al.
Veröffentlicht: (2024)
von: Vijayvergiya, Manushree, et al.
Veröffentlicht: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
von: Jha, Smriti, et al.
Veröffentlicht: (2026)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study
von: Vella, Annie, et al.
Veröffentlicht: (2026)
von: Vella, Annie, et al.
Veröffentlicht: (2026)
Lessons from Building StackSpot AI: A Contextualized AI Coding Assistant
von: Pinto, Gustavo, et al.
Veröffentlicht: (2023)
von: Pinto, Gustavo, et al.
Veröffentlicht: (2023)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
von: Chen, Valerie, et al.
Veröffentlicht: (2026)
von: Chen, Valerie, et al.
Veröffentlicht: (2026)
Assessing AI-Based Code Assistants in Method Generation Tasks
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
Harnessing Hype to Teach Empirical Thinking: An Experience With AI Coding Assistants
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
von: Wyrich, Marvin, et al.
Veröffentlicht: (2026)
Usage, Effects and Requirements for AI Coding Assistants in the Enterprise: An Empirical Study
von: Vukovic, Maja, et al.
Veröffentlicht: (2026)
von: Vukovic, Maja, et al.
Veröffentlicht: (2026)
Harnessing the Potential of Gen-AI Coding Assistants in Public Sector Software Development
von: Ng, Kevin KB, et al.
Veröffentlicht: (2024)
von: Ng, Kevin KB, et al.
Veröffentlicht: (2024)
Generating Java Methods: An Empirical Assessment of Four AI-Based Code Assistants
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
von: Corso, Vincenzo, et al.
Veröffentlicht: (2024)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoyu, et al.
Veröffentlicht: (2025)
Inducing Vulnerable Code Generation in LLM Coding Assistants
von: Zeng, Binqi, et al.
Veröffentlicht: (2025)
von: Zeng, Binqi, et al.
Veröffentlicht: (2025)
Does Co-Development with AI Assistants Lead to More Maintainable Code? A Registered Report
von: Borg, Markus, et al.
Veröffentlicht: (2024)
von: Borg, Markus, et al.
Veröffentlicht: (2024)
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
von: Adnan, Bassam, et al.
Veröffentlicht: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
von: Liu, Linbo, et al.
Veröffentlicht: (2025)
von: Liu, Linbo, et al.
Veröffentlicht: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
von: Guo, Hanyang, et al.
Veröffentlicht: (2025)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
von: Li, Xuefeng, et al.
Veröffentlicht: (2026)
von: Li, Xuefeng, et al.
Veröffentlicht: (2026)
Collaborator or Assistant? How AI Coding Agents Partition Work Across Pull Request Lifecycles
von: Jo, Young, et al.
Veröffentlicht: (2026)
von: Jo, Young, et al.
Veröffentlicht: (2026)
Spec-Driven Development:From Code to Contract in the Age of AI Coding Assistants
von: Piskala, Deepak Babu
Veröffentlicht: (2026)
von: Piskala, Deepak Babu
Veröffentlicht: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
How is Google using AI for internal code migrations?
von: Nikolov, Stoyan, et al.
Veröffentlicht: (2025)
von: Nikolov, Stoyan, et al.
Veröffentlicht: (2025)
Towards Verified Code Reasoning by LLMs
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
von: Sistla, Meghana, et al.
Veröffentlicht: (2025)
How Agentic AI Coding Assistants Become the Attacker's Shell
von: Liu, Yue, et al.
Veröffentlicht: (2026)
von: Liu, Yue, et al.
Veröffentlicht: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
VecIntrinBench: Benchmarking Cross-Architecture Intrinsic Code Migration for RISC-V Vector
von: Han, Liutong, et al.
Veröffentlicht: (2025)
von: Han, Liutong, et al.
Veröffentlicht: (2025)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CRQBench: A Benchmark of Code Reasoning Questions
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2024) -
On the Effectiveness of Modular Testing in EvoSuite
von: Dinella, Elizabeth
Veröffentlicht: (2026) -
Program Structure Aware Precondition Generation
von: Dinella, Elizabeth, et al.
Veröffentlicht: (2023) -
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
von: Mathai, Alex, et al.
Veröffentlicht: (2024) -
Agentic Code Reasoning
von: Ugare, Shubham, et al.
Veröffentlicht: (2026)