Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shehada, Kareem, Wu, Yifan, Feng, Wyatt D., Iyer, Adithya, Kumfert, Gryphon, Ding, Yangruibo, Qian, Zhiyun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911276770263040
author Shehada, Kareem
Wu, Yifan
Feng, Wyatt D.
Iyer, Adithya
Kumfert, Gryphon
Ding, Yangruibo
Qian, Zhiyun
author_facet Shehada, Kareem
Wu, Yifan
Feng, Wyatt D.
Iyer, Adithya
Kumfert, Gryphon
Ding, Yangruibo
Qian, Zhiyun
contents Large Language Models (LLMs) have revolutionized automated program repair (APR) but current benchmarks like SWE-Bench predominantly focus on userspace applications and overlook the complexities of kernel-space debugging and repair. The Linux kernel poses unique challenges due to its monolithic structure, concurrency, and low-level hardware interactions. Prior efforts such as KGym and CrashFixer have highlighted the difficulty of APR in this domain, reporting low success rates or relying on costly and complex pipelines and pricey cloud infrastructure. In this work, we introduce RGym, a lightweight, platform-agnostic APR evaluation framework for the Linux kernel designed to operate on local commodity hardware. Built on RGym, we propose a simple yet effective APR pipeline leveraging specialized localization techniques (e.g., call stacks and blamed commits) to overcome the unrealistic usage of oracles in KGym. We test on a filtered and verified dataset of 143 bugs. Our method achieves up to a 43.36% pass rate with GPT-5 Thinking while maintaining a cost of under $0.20 per bug. We further conduct an ablation study to analyze contributions from our proposed localization strategy, prompt structure, and model choice, and demonstrate that feedback-based retries can significantly enhance success rates.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
Shehada, Kareem
Wu, Yifan
Feng, Wyatt D.
Iyer, Adithya
Kumfert, Gryphon
Ding, Yangruibo
Qian, Zhiyun
Software Engineering
Large Language Models (LLMs) have revolutionized automated program repair (APR) but current benchmarks like SWE-Bench predominantly focus on userspace applications and overlook the complexities of kernel-space debugging and repair. The Linux kernel poses unique challenges due to its monolithic structure, concurrency, and low-level hardware interactions. Prior efforts such as KGym and CrashFixer have highlighted the difficulty of APR in this domain, reporting low success rates or relying on costly and complex pipelines and pricey cloud infrastructure. In this work, we introduce RGym, a lightweight, platform-agnostic APR evaluation framework for the Linux kernel designed to operate on local commodity hardware. Built on RGym, we propose a simple yet effective APR pipeline leveraging specialized localization techniques (e.g., call stacks and blamed commits) to overcome the unrealistic usage of oracles in KGym. We test on a filtered and verified dataset of 143 bugs. Our method achieves up to a 43.36% pass rate with GPT-5 Thinking while maintaining a cost of under $0.20 per bug. We further conduct an ablation study to analyze contributions from our proposed localization strategy, prompt structure, and model choice, and demonstrate that feedback-based retries can significantly enhance success rates.
title Rethinking Kernel Program Repair: Benchmarking and Enhancing LLMs with RGym
topic Software Engineering
url https://arxiv.org/abs/2511.15757