Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zhenhao, Huang, Zhuochen, He, Yike, Wang, Chong, Wang, Jiajun, Wu, Yijian, Peng, Xin, Lou, Yiling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914622333779968
author Zhou, Zhenhao
Huang, Zhuochen
He, Yike
Wang, Chong
Wang, Jiajun
Wu, Yijian
Peng, Xin
Lou, Yiling
author_facet Zhou, Zhenhao
Huang, Zhuochen
He, Yike
Wang, Chong
Wang, Jiajun
Wu, Yijian
Peng, Xin
Lou, Yiling
contents The Linux kernel is a critical system, serving as the foundation for numerous systems. Bugs in the Linux kernel can cause serious consequences, affecting billions of users. Fault localization (FL), which aims at identifying the buggy code elements in software, plays an essential role in software quality assurance. While recent LLM agents have achieved promising accuracy in FL on recent benchmarks like SWE-bench, it remains unclear how well these methods perform in the Linux kernel, where FL is much more challenging due to the large-scale code base, limited observability, and diverse impact factors. In this paper, we introduce LinuxFLBench, a FL benchmark constructed from real-world Linux kernel bugs. We conduct an empirical study to assess the performance of state-of-the-art LLM agents on the Linux kernel. Our initial results reveal that existing agents struggle with this task, achieving a best top-1 accuracy of only 41.6% at file level. To address this challenge, we propose LinuxFL$^+$, an enhancement framework designed to improve FL effectiveness of LLM agents for the Linux kernel. LinuxFL$^+$ substantially improves the FL accuracy of all studied agents (e.g., 7.2% - 11.2% accuracy increase) with minimal costs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19489
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
Zhou, Zhenhao
Huang, Zhuochen
He, Yike
Wang, Chong
Wang, Jiajun
Wu, Yijian
Peng, Xin
Lou, Yiling
Artificial Intelligence
Software Engineering
The Linux kernel is a critical system, serving as the foundation for numerous systems. Bugs in the Linux kernel can cause serious consequences, affecting billions of users. Fault localization (FL), which aims at identifying the buggy code elements in software, plays an essential role in software quality assurance. While recent LLM agents have achieved promising accuracy in FL on recent benchmarks like SWE-bench, it remains unclear how well these methods perform in the Linux kernel, where FL is much more challenging due to the large-scale code base, limited observability, and diverse impact factors. In this paper, we introduce LinuxFLBench, a FL benchmark constructed from real-world Linux kernel bugs. We conduct an empirical study to assess the performance of state-of-the-art LLM agents on the Linux kernel. Our initial results reveal that existing agents struggle with this task, achieving a best top-1 accuracy of only 41.6% at file level. To address this challenge, we propose LinuxFL$^+$, an enhancement framework designed to improve FL effectiveness of LLM agents for the Linux kernel. LinuxFL$^+$ substantially improves the FL accuracy of all studied agents (e.g., 7.2% - 11.2% accuracy increase) with minimal costs.
title Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2505.19489