LLPut: Investigating Large Language Models for Bug Report-Based Input Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hasan, Alif Al, Saha, Subarna, Imran, Mia Mohammad, Zaman, Tarannum Shaila
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915673625591808
author Hasan, Alif Al
Saha, Subarna
Imran, Mia Mohammad
Zaman, Tarannum Shaila
author_facet Hasan, Alif Al
Saha, Subarna
Imran, Mia Mohammad
Zaman, Tarannum Shaila
contents Failure-inducing inputs play a crucial role in diagnosing and analyzing software bugs. Bug reports typically contain these inputs, which developers extract to facilitate debugging. Since bug reports are written in natural language, prior research has leveraged various Natural Language Processing (NLP) techniques for automated input extraction. With the advent of Large Language Models (LLMs), an important research question arises: how effectively can generative LLMs extract failure-inducing inputs from bug reports? In this paper, we propose LLPut, a technique to empirically evaluate the performance of three open-source generative LLMs -- LLaMA, Qwen, and Qwen-Coder -- in extracting relevant inputs from bug reports. We conduct an experimental evaluation on a dataset of 206 bug reports to assess the accuracy and effectiveness of these models. Our findings provide insights into the capabilities and limitations of generative LLMs in automated bug diagnosis.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20578
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLPut: Investigating Large Language Models for Bug Report-Based Input Generation
Hasan, Alif Al
Saha, Subarna
Imran, Mia Mohammad
Zaman, Tarannum Shaila
Software Engineering
Failure-inducing inputs play a crucial role in diagnosing and analyzing software bugs. Bug reports typically contain these inputs, which developers extract to facilitate debugging. Since bug reports are written in natural language, prior research has leveraged various Natural Language Processing (NLP) techniques for automated input extraction. With the advent of Large Language Models (LLMs), an important research question arises: how effectively can generative LLMs extract failure-inducing inputs from bug reports? In this paper, we propose LLPut, a technique to empirically evaluate the performance of three open-source generative LLMs -- LLaMA, Qwen, and Qwen-Coder -- in extracting relevant inputs from bug reports. We conduct an experimental evaluation on a dataset of 206 bug reports to assess the accuracy and effectiveness of these models. Our findings provide insights into the capabilities and limitations of generative LLMs in automated bug diagnosis.
title LLPut: Investigating Large Language Models for Bug Report-Based Input Generation
topic Software Engineering
url https://arxiv.org/abs/2503.20578