StackEval: Benchmarking LLMs in Coding Assistance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Nidhish, Genc, Zulkuf, Araci, Dogu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929619558465536
author Shah, Nidhish
Genc, Zulkuf
Araci, Dogu
author_facet Shah, Nidhish
Genc, Zulkuf
Araci, Dogu
contents We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack Overflow questions, and StackUnseen, a dynamic benchmark featuring the most recent Stack Overflow content. These benchmarks offer novel insights into the capabilities and limitations of LLMs, particularly in handling new and emerging content. Additionally, we assess LLMs' proficiency as judges for coding tasks using a curated, human-annotated dataset, exploring their evaluation capabilities and potential biases, including whether they favor their own generated solutions. Our findings underscore the potential of these benchmarks to advance LLM development and application in coding assistance. To ensure reproducibility, we publicly share our datasets and evaluation code at https://github.com/ProsusAI/stack-eval .
format Preprint
id arxiv_https___arxiv_org_abs_2412_05288
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StackEval: Benchmarking LLMs in Coding Assistance
Shah, Nidhish
Genc, Zulkuf
Araci, Dogu
Software Engineering
Computation and Language
Machine Learning
We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack Overflow questions, and StackUnseen, a dynamic benchmark featuring the most recent Stack Overflow content. These benchmarks offer novel insights into the capabilities and limitations of LLMs, particularly in handling new and emerging content. Additionally, we assess LLMs' proficiency as judges for coding tasks using a curated, human-annotated dataset, exploring their evaluation capabilities and potential biases, including whether they favor their own generated solutions. Our findings underscore the potential of these benchmarks to advance LLM development and application in coding assistance. To ensure reproducibility, we publicly share our datasets and evaluation code at https://github.com/ProsusAI/stack-eval .
title StackEval: Benchmarking LLMs in Coding Assistance
topic Software Engineering
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.05288