Saved in:
| Main Authors: | Mathai, Alex, Huang, Chenxi, Maniatis, Petros, Nogikh, Aleksandr, Ivancic, Franjo, Yang, Junfeng, Ray, Baishakhi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.02680 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026)
by: Huang, Chenxi, et al.
Published: (2026)
CrashFixer: A crash resolution agent for the Linux kernel
by: Mathai, Alex, et al.
Published: (2025)
by: Mathai, Alex, et al.
Published: (2025)
SemAgent: A Semantics Aware Program Repair Agent
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
CRQBench: A Benchmark of Code Reasoning Questions
by: Dinella, Elizabeth, et al.
Published: (2024)
by: Dinella, Elizabeth, et al.
Published: (2024)
Beyond Crash-to-Patch: Patch Evolution for Linux Kernel Repair
by: Bai, Luyao, et al.
Published: (2026)
by: Bai, Luyao, et al.
Published: (2026)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
RubberDuckBench: A Benchmark for AI Coding Assistants
by: Mohammed, Ferida, et al.
Published: (2026)
by: Mohammed, Ferida, et al.
Published: (2026)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
by: Shi, Sherry, et al.
Published: (2025)
by: Shi, Sherry, et al.
Published: (2025)
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
by: Nitin, Vikram, et al.
Published: (2024)
by: Nitin, Vikram, et al.
Published: (2024)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025)
by: Kim, Myeongsoo, et al.
Published: (2025)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
by: Peng, Jinjun, et al.
Published: (2025)
by: Peng, Jinjun, et al.
Published: (2025)
Rationale Dataset and Analysis for the Commit Messages of the Linux Kernel Out-of-Memory Killer
by: Dhaouadi, Mouna, et al.
Published: (2024)
by: Dhaouadi, Mouna, et al.
Published: (2024)
Yuga: Automatically Detecting Lifetime Annotation Bugs in the Rust Language
by: Nitin, Vikram, et al.
Published: (2023)
by: Nitin, Vikram, et al.
Published: (2023)
Rusty Linux: Advances in Rust for Linux Kernel Development
by: Panter, Shane K., et al.
Published: (2024)
by: Panter, Shane K., et al.
Published: (2024)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
Linux Kernel Configurations at Scale: A Dataset for Performance and Evolution Analysis
by: Borges, Heraldo, et al.
Published: (2025)
by: Borges, Heraldo, et al.
Published: (2025)
CodeSense: a Real-World Benchmark and Dataset for Code Semantic Reasoning
by: Roy, Monoshi Kumar, et al.
Published: (2025)
by: Roy, Monoshi Kumar, et al.
Published: (2025)
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
A Time Series Analysis of Assertions in the Linux Kernel
by: Ruohonen, Jukka
Published: (2024)
by: Ruohonen, Jukka
Published: (2024)
Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel
by: Lyu, Yunbo, et al.
Published: (2023)
by: Lyu, Yunbo, et al.
Published: (2023)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Trustworthy AI Software Engineers
by: Aleti, Aldeida, et al.
Published: (2026)
by: Aleti, Aldeida, et al.
Published: (2026)
Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Characterizing and Mitigating False-Positive Bug Reports in the Linux Kernel
by: Tian, Jiashuo, et al.
Published: (2026)
by: Tian, Jiashuo, et al.
Published: (2026)
CrashJS: A NodeJS Benchmark for Automated Crash Reproduction
by: Oliver, Philip, et al.
Published: (2024)
by: Oliver, Philip, et al.
Published: (2024)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Automatic Programming: Large Language Models and Beyond
by: Lyu, Michael R., et al.
Published: (2024)
by: Lyu, Michael R., et al.
Published: (2024)
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
by: Wang, Yiran, et al.
Published: (2025)
by: Wang, Yiran, et al.
Published: (2025)
EditLord: Learning Code Transformation Rules for Code Editing
by: Li, Weichen, et al.
Published: (2025)
by: Li, Weichen, et al.
Published: (2025)
An Investigation of Patch Porting Practices of the Linux Kernel Ecosystem
by: Li, Xingyu, et al.
Published: (2024)
by: Li, Xingyu, et al.
Published: (2024)
C2SaferRust: Transforming C Projects into Safer Rust with NeuroSymbolic Techniques
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
Crash Report Enhancement with Large Language Models: An Empirical Study
by: Fahim, S M Farah Al, et al.
Published: (2025)
by: Fahim, S M Farah Al, et al.
Published: (2025)
MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
by: Dang, Pucheng, et al.
Published: (2025)
by: Dang, Pucheng, et al.
Published: (2025)
Agentic AI Software Engineers: Programming with Trust
by: Roychoudhury, Abhik, et al.
Published: (2025)
by: Roychoudhury, Abhik, et al.
Published: (2025)
Automatic Detection of Reference Counting Bugs in Linux Kernel Drivers
by: Hattori, Joe, et al.
Published: (2026)
by: Hattori, Joe, et al.
Published: (2026)
LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
by: Kharlamova, Arina, et al.
Published: (2025)
by: Kharlamova, Arina, et al.
Published: (2025)
Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
by: Ceka, Ira, et al.
Published: (2025)
by: Ceka, Ira, et al.
Published: (2025)
ScarfBench: A Benchmark for Cross-Framework Application Migration in Enterprise Java
by: Pavuluri, Advait, et al.
Published: (2026)
by: Pavuluri, Advait, et al.
Published: (2026)
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
by: Zhou, Zhenhao, et al.
Published: (2025)
by: Zhou, Zhenhao, et al.
Published: (2025)
Similar Items
-
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026) -
CrashFixer: A crash resolution agent for the Linux kernel
by: Mathai, Alex, et al.
Published: (2025) -
SemAgent: A Semantics Aware Program Repair Agent
by: Pabba, Anvith, et al.
Published: (2025) -
CRQBench: A Benchmark of Code Reasoning Questions
by: Dinella, Elizabeth, et al.
Published: (2024) -
Beyond Crash-to-Patch: Patch Evolution for Linux Kernel Repair
by: Bai, Luyao, et al.
Published: (2026)