LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kuhar, Sachit, Ahmad, Wasi Uddin, Wang, Zijian, Jain, Nihal, Qian, Haifeng, Ray, Baishakhi, Ramanathan, Murali Krishna, Ma, Xiaofei, Deoras, Anoop
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908351025119232
author Kuhar, Sachit
Ahmad, Wasi Uddin
Wang, Zijian
Jain, Nihal
Qian, Haifeng
Ray, Baishakhi
Ramanathan, Murali Krishna
Ma, Xiaofei
Deoras, Anoop
author_facet Kuhar, Sachit
Ahmad, Wasi Uddin
Wang, Zijian
Jain, Nihal
Qian, Haifeng
Ray, Baishakhi
Ramanathan, Murali Krishna
Ma, Xiaofei
Deoras, Anoop
contents Recent advancements in code completion models have primarily focused on local file contexts. However, these studies do not fully capture the complexity of real-world software development, which often requires the use of rapidly-evolving public libraries. To fill the gap, we introduce LibEvolutionEval, a detailed study requiring an understanding of library evolution to perform in-line code completion accurately. LibEvolutionEval provides a version-specific code-completion task comprised of eight libraries (torch, torchvision, scipy, pil, tqdm, pyyaml, matplotlib, and pandas) as they evolve over the year along with a detailed analysis of the evolution of two popular and well-maintained public libraries: PyTorch and Matplotlib. We evaluate popular public models and find that public library evolution significantly influences model performance. We explored mitigation methods by studying how retrieved version-specific library documentation and prompting can improve the model's capability in handling these fast-evolving packages, paving a promising future path in better handling fast-evolving libraries.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04478
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
Kuhar, Sachit
Ahmad, Wasi Uddin
Wang, Zijian
Jain, Nihal
Qian, Haifeng
Ray, Baishakhi
Ramanathan, Murali Krishna
Ma, Xiaofei
Deoras, Anoop
Software Engineering
Artificial Intelligence
Recent advancements in code completion models have primarily focused on local file contexts. However, these studies do not fully capture the complexity of real-world software development, which often requires the use of rapidly-evolving public libraries. To fill the gap, we introduce LibEvolutionEval, a detailed study requiring an understanding of library evolution to perform in-line code completion accurately. LibEvolutionEval provides a version-specific code-completion task comprised of eight libraries (torch, torchvision, scipy, pil, tqdm, pyyaml, matplotlib, and pandas) as they evolve over the year along with a detailed analysis of the evolution of two popular and well-maintained public libraries: PyTorch and Matplotlib. We evaluate popular public models and find that public library evolution significantly influences model performance. We explored mitigation methods by studying how retrieved version-specific library documentation and prompting can improve the model's capability in handling these fast-evolving packages, paving a promising future path in better handling fast-evolving libraries.
title LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2412.04478