StackEval: Benchmarking LLMs in Coding Assistance
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929619558465536 |
|---|---|
| author | Shah, Nidhish Genc, Zulkuf Araci, Dogu |
| author_facet | Shah, Nidhish Genc, Zulkuf Araci, Dogu |
| contents | We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack Overflow questions, and StackUnseen, a dynamic benchmark featuring the most recent Stack Overflow content. These benchmarks offer novel insights into the capabilities and limitations of LLMs, particularly in handling new and emerging content. Additionally, we assess LLMs' proficiency as judges for coding tasks using a curated, human-annotated dataset, exploring their evaluation capabilities and potential biases, including whether they favor their own generated solutions. Our findings underscore the potential of these benchmarks to advance LLM development and application in coding assistance. To ensure reproducibility, we publicly share our datasets and evaluation code at https://github.com/ProsusAI/stack-eval . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_05288 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | StackEval: Benchmarking LLMs in Coding Assistance Shah, Nidhish Genc, Zulkuf Araci, Dogu Software Engineering Computation and Language Machine Learning We present two comprehensive benchmarks to evaluate the performance of language models in coding assistance tasks, covering code writing, debugging, code review, and conceptual understanding. Our main contribution includes two curated datasets: StackEval, a large-scale benchmark derived from Stack Overflow questions, and StackUnseen, a dynamic benchmark featuring the most recent Stack Overflow content. These benchmarks offer novel insights into the capabilities and limitations of LLMs, particularly in handling new and emerging content. Additionally, we assess LLMs' proficiency as judges for coding tasks using a curated, human-annotated dataset, exploring their evaluation capabilities and potential biases, including whether they favor their own generated solutions. Our findings underscore the potential of these benchmarks to advance LLM development and application in coding assistance. To ensure reproducibility, we publicly share our datasets and evaluation code at https://github.com/ProsusAI/stack-eval . |
| title | StackEval: Benchmarking LLMs in Coding Assistance |
| topic | Software Engineering Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2412.05288 |