Crash Report Enhancement with Large Language Models: An Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fahim, S M Farah Al, Rafi, Md Nakhla, Ma, Zeyang, Kim, Dong Jae, Tse-Hsun, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914041395412992
author Fahim, S M Farah Al
Rafi, Md Nakhla
Ma, Zeyang
Kim, Dong Jae
Tse-Hsun
Chen
author_facet Fahim, S M Farah Al
Rafi, Md Nakhla
Ma, Zeyang
Kim, Dong Jae
Tse-Hsun
Chen
contents Crash reports are central to software maintenance, yet many lack the diagnostic detail developers need to debug efficiently. We examine whether large language models can enhance crash reports by adding fault locations, root-cause explanations, and repair suggestions. We study two enhancement strategies: Direct-LLM, a single-shot approach that uses stack-trace context, and Agentic-LLM, an iterative approach that explores the repository for additional evidence. On a dataset of 492 real-world crash reports, LLM-enhanced reports improve Top-1 problem-localization accuracy from 10.6% (original reports) to 40.2-43.1%, and produce suggested fixes that closely resemble developer patches (CodeBLEU around 56-57%). Both our manual evaluations and LLM-as-a-judge assessment show that Agentic-LLM delivers stronger root-cause explanations and more actionable repair guidance. A user study with 16 participants further confirms that enhanced reports make crashes easier to understand and resolve, with the largest improvement in repair guidance. These results indicate that supplying LLMs with stack traces and repository code yields enhanced crash reports that are substantially more useful for debugging.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13535
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Crash Report Enhancement with Large Language Models: An Empirical Study
Fahim, S M Farah Al
Rafi, Md Nakhla
Ma, Zeyang
Kim, Dong Jae
Tse-Hsun
Chen
Software Engineering
Crash reports are central to software maintenance, yet many lack the diagnostic detail developers need to debug efficiently. We examine whether large language models can enhance crash reports by adding fault locations, root-cause explanations, and repair suggestions. We study two enhancement strategies: Direct-LLM, a single-shot approach that uses stack-trace context, and Agentic-LLM, an iterative approach that explores the repository for additional evidence. On a dataset of 492 real-world crash reports, LLM-enhanced reports improve Top-1 problem-localization accuracy from 10.6% (original reports) to 40.2-43.1%, and produce suggested fixes that closely resemble developer patches (CodeBLEU around 56-57%). Both our manual evaluations and LLM-as-a-judge assessment show that Agentic-LLM delivers stronger root-cause explanations and more actionable repair guidance. A user study with 16 participants further confirms that enhanced reports make crashes easier to understand and resolve, with the largest improvement in repair guidance. These results indicate that supplying LLMs with stack traces and repository code yields enhanced crash reports that are substantially more useful for debugging.
title Crash Report Enhancement with Large Language Models: An Empirical Study
topic Software Engineering
url https://arxiv.org/abs/2509.13535