KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
Fuente:
arXiv
Salvato in:
| Autori principali: | Mathai, Alex, Huang, Chenxi, Maniatis, Petros, Nogikh, Aleksandr, Ivancic, Franjo, Yang, Junfeng, Ray, Baishakhi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
di: Huang, Chenxi, et al.
Pubblicazione: (2026)
di: Huang, Chenxi, et al.
Pubblicazione: (2026)
CrashFixer: A crash resolution agent for the Linux kernel
di: Mathai, Alex, et al.
Pubblicazione: (2025)
di: Mathai, Alex, et al.
Pubblicazione: (2025)
Beyond Crash-to-Patch: Patch Evolution for Linux Kernel Repair
di: Bai, Luyao, et al.
Pubblicazione: (2026)
di: Bai, Luyao, et al.
Pubblicazione: (2026)
CRQBench: A Benchmark of Code Reasoning Questions
di: Dinella, Elizabeth, et al.
Pubblicazione: (2024)
di: Dinella, Elizabeth, et al.
Pubblicazione: (2024)
SemAgent: A Semantics Aware Program Repair Agent
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
di: Pabba, Anvith, et al.
Pubblicazione: (2025)
RubberDuckBench: A Benchmark for AI Coding Assistants
di: Mohammed, Ferida, et al.
Pubblicazione: (2026)
di: Mohammed, Ferida, et al.
Pubblicazione: (2026)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
di: Chen, Simin, et al.
Pubblicazione: (2025)
di: Chen, Simin, et al.
Pubblicazione: (2025)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
di: Shi, Sherry, et al.
Pubblicazione: (2025)
di: Shi, Sherry, et al.
Pubblicazione: (2025)
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
di: Nitin, Vikram, et al.
Pubblicazione: (2024)
di: Nitin, Vikram, et al.
Pubblicazione: (2024)
Rationale Dataset and Analysis for the Commit Messages of the Linux Kernel Out-of-Memory Killer
di: Dhaouadi, Mouna, et al.
Pubblicazione: (2024)
di: Dhaouadi, Mouna, et al.
Pubblicazione: (2024)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
di: Kim, Myeongsoo, et al.
Pubblicazione: (2025)
di: Kim, Myeongsoo, et al.
Pubblicazione: (2025)
Yuga: Automatically Detecting Lifetime Annotation Bugs in the Rust Language
di: Nitin, Vikram, et al.
Pubblicazione: (2023)
di: Nitin, Vikram, et al.
Pubblicazione: (2023)
Rusty Linux: Advances in Rust for Linux Kernel Development
di: Panter, Shane K., et al.
Pubblicazione: (2024)
di: Panter, Shane K., et al.
Pubblicazione: (2024)
Linux Kernel Configurations at Scale: A Dataset for Performance and Evolution Analysis
di: Borges, Heraldo, et al.
Pubblicazione: (2025)
di: Borges, Heraldo, et al.
Pubblicazione: (2025)
A Time Series Analysis of Assertions in the Linux Kernel
di: Ruohonen, Jukka
Pubblicazione: (2024)
di: Ruohonen, Jukka
Pubblicazione: (2024)
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel
di: Lyu, Yunbo, et al.
Pubblicazione: (2023)
di: Lyu, Yunbo, et al.
Pubblicazione: (2023)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
di: Peng, Jinjun, et al.
Pubblicazione: (2025)
di: Peng, Jinjun, et al.
Pubblicazione: (2025)
CrashJS: A NodeJS Benchmark for Automated Crash Reproduction
di: Oliver, Philip, et al.
Pubblicazione: (2024)
di: Oliver, Philip, et al.
Pubblicazione: (2024)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
di: Roy, Monoshi Kumar, et al.
Pubblicazione: (2025)
di: Roy, Monoshi Kumar, et al.
Pubblicazione: (2025)
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
di: Tian, Jiashuo, et al.
Pubblicazione: (2026)
di: Tian, Jiashuo, et al.
Pubblicazione: (2026)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
di: Cheng, Runxiang, et al.
Pubblicazione: (2026)
di: Cheng, Runxiang, et al.
Pubblicazione: (2026)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
di: Wang, Yiran, et al.
Pubblicazione: (2025)
di: Wang, Yiran, et al.
Pubblicazione: (2025)
Trustworthy AI Software Engineers
di: Aleti, Aldeida, et al.
Pubblicazione: (2026)
di: Aleti, Aldeida, et al.
Pubblicazione: (2026)
Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities
di: Chen, Simin, et al.
Pubblicazione: (2025)
di: Chen, Simin, et al.
Pubblicazione: (2025)
Crash Report Enhancement with Large Language Models: An Empirical Study
di: Fahim, S M Farah Al, et al.
Pubblicazione: (2025)
di: Fahim, S M Farah Al, et al.
Pubblicazione: (2025)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
An Investigation of Patch Porting Practices of the Linux Kernel Ecosystem
di: Li, Xingyu, et al.
Pubblicazione: (2024)
di: Li, Xingyu, et al.
Pubblicazione: (2024)
Automatic Programming: Large Language Models and Beyond
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
di: Lyu, Michael R., et al.
Pubblicazione: (2024)
C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
di: Nitin, Vikram, et al.
Pubblicazione: (2025)
Automatic Detection of Reference Counting Bugs in Linux Kernel Drivers
di: Hattori, Joe, et al.
Pubblicazione: (2026)
di: Hattori, Joe, et al.
Pubblicazione: (2026)
LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
di: Kharlamova, Arina, et al.
Pubblicazione: (2025)
di: Kharlamova, Arina, et al.
Pubblicazione: (2025)
Fast Fixes and Faulty Drivers: An Empirical Analysis of Regression Bug Fixing Times in the Linux Kernel
di: Ruohonen, Jukka, et al.
Pubblicazione: (2024)
di: Ruohonen, Jukka, et al.
Pubblicazione: (2024)
MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
di: Dang, Pucheng, et al.
Pubblicazione: (2025)
di: Dang, Pucheng, et al.
Pubblicazione: (2025)
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
di: Zhou, Zhenhao, et al.
Pubblicazione: (2025)
di: Zhou, Zhenhao, et al.
Pubblicazione: (2025)
EditLord: Learning Code Transformation Rules for Code Editing
di: Li, Weichen, et al.
Pubblicazione: (2025)
di: Li, Weichen, et al.
Pubblicazione: (2025)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
di: Chen, Simin, et al.
Pubblicazione: (2025)
di: Chen, Simin, et al.
Pubblicazione: (2025)
ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java
di: Pavuluri, Advait, et al.
Pubblicazione: (2026)
di: Pavuluri, Advait, et al.
Pubblicazione: (2026)
Linux Kernel Recency Matters, CVE Severity Doesn't, and History Fades
di: Przymus, Piotr, et al.
Pubblicazione: (2026)
di: Przymus, Piotr, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
di: Huang, Chenxi, et al.
Pubblicazione: (2026) -
CrashFixer: A crash resolution agent for the Linux kernel
di: Mathai, Alex, et al.
Pubblicazione: (2025) -
Beyond Crash-to-Patch: Patch Evolution for Linux Kernel Repair
di: Bai, Luyao, et al.
Pubblicazione: (2026) -
CRQBench: A Benchmark of Code Reasoning Questions
di: Dinella, Elizabeth, et al.
Pubblicazione: (2024) -
SemAgent: A Semantics Aware Program Repair Agent
di: Pabba, Anvith, et al.
Pubblicazione: (2025)