Assessing the Answerability of Queries in Retrieval-Augmented Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Geonmin, Kim, Jaeyeon, Park, Hancheol, Shin, Wooksu, Kim, Tae-Ho
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915032116232192
author Kim, Geonmin
Kim, Jaeyeon
Park, Hancheol
Shin, Wooksu
Kim, Tae-Ho
author_facet Kim, Geonmin
Kim, Jaeyeon
Park, Hancheol
Shin, Wooksu
Kim, Tae-Ho
contents Thanks to unprecedented language understanding and generation capabilities of large language model (LLM), Retrieval-augmented Code Generation (RaCG) has recently been widely utilized among software developers. While this has increased productivity, there are still frequent instances of incorrect codes being provided. In particular, there are cases where plausible yet incorrect codes are generated for queries from users that cannot be answered with the given queries and API descriptions. This study proposes a task for evaluating answerability, which assesses whether valid answers can be generated based on users' queries and retrieved APIs in RaCG. Additionally, we build a benchmark dataset called Retrieval-augmented Code Generability Evaluation (RaCGEval) to evaluate the performance of models performing this task. Experimental results show that this task remains at a very challenging level, with baseline models exhibiting a low performance of 46.7%. Furthermore, this study discusses methods that could significantly improve performance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing the Answerability of Queries in Retrieval-Augmented Code Generation
Kim, Geonmin
Kim, Jaeyeon
Park, Hancheol
Shin, Wooksu
Kim, Tae-Ho
Computation and Language
Thanks to unprecedented language understanding and generation capabilities of large language model (LLM), Retrieval-augmented Code Generation (RaCG) has recently been widely utilized among software developers. While this has increased productivity, there are still frequent instances of incorrect codes being provided. In particular, there are cases where plausible yet incorrect codes are generated for queries from users that cannot be answered with the given queries and API descriptions. This study proposes a task for evaluating answerability, which assesses whether valid answers can be generated based on users' queries and retrieved APIs in RaCG. Additionally, we build a benchmark dataset called Retrieval-augmented Code Generability Evaluation (RaCGEval) to evaluate the performance of models performing this task. Experimental results show that this task remains at a very challenging level, with baseline models exhibiting a low performance of 46.7%. Furthermore, this study discusses methods that could significantly improve performance.
title Assessing the Answerability of Queries in Retrieval-Augmented Code Generation
topic Computation and Language
url https://arxiv.org/abs/2411.05547