AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Tanghaoran, Mao, Xinjun, Wang, Shangwen, Zhao, Yuxin, Lu, Yao, Zhang, Jin, Zhang, Zhang, Yang, Kang, Yu, Yue
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909984399294464
author Zhang, Tanghaoran
Mao, Xinjun
Wang, Shangwen
Zhao, Yuxin
Lu, Yao
Zhang, Jin
Zhang, Zhang
Yang, Kang
Yu, Yue
author_facet Zhang, Tanghaoran
Mao, Xinjun
Wang, Shangwen
Zhao, Yuxin
Lu, Yao
Zhang, Jin
Zhang, Zhang
Yang, Kang
Yu, Yue
contents Recent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs' performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, Practical Context. Tasks in AdaptEval are derived from developers' practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, Multi-granularity Annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, Fine-grained Evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs' performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs' adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs' capabilities in code snippet adaptation, supporting their real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04540
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
Zhang, Tanghaoran
Mao, Xinjun
Wang, Shangwen
Zhao, Yuxin
Lu, Yao
Zhang, Jin
Zhang, Zhang
Yang, Kang
Yu, Yue
Software Engineering
Artificial Intelligence
Recent advancements in large language models (LLMs) have automated various software engineering tasks, with benchmarks emerging to evaluate their capabilities. However, for adaptation, a critical activity during code reuse, there is no benchmark to assess LLMs' performance, leaving their practical utility in this area unclear. To fill this gap, we propose AdaptEval, a benchmark designed to evaluate LLMs on code snippet adaptation. Unlike existing benchmarks, AdaptEval incorporates the following three distinctive features: First, Practical Context. Tasks in AdaptEval are derived from developers' practices, preserving rich contextual information from Stack Overflow and GitHub communities. Second, Multi-granularity Annotation. Each task is annotated with requirements at both task and adaptation levels, supporting the evaluation of LLMs across diverse adaptation scenarios. Third, Fine-grained Evaluation. AdaptEval includes a two-tier testing framework combining adaptation-level and function-level tests, which enables evaluating LLMs' performance across various individual adaptations. Based on AdaptEval, we conduct the first empirical study to evaluate six instruction-tuned LLMs and especially three reasoning LLMs on code snippet adaptation. Experimental results demonstrate that AdaptEval enables the assessment of LLMs' adaptation capabilities from various perspectives. It also provides critical insights into their current limitations, particularly their struggle to follow explicit instructions. We hope AdaptEval can facilitate further investigation and enhancement of LLMs' capabilities in code snippet adaptation, supporting their real-world applications.
title AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2601.04540