Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Chenxi, Mathai, Alex, Yu, Feiyang, Nogikh, Aleksandr, Maniatis, Petros, Ivančić, Franjo, Wu, Eugene, Kaffes, Kostis, Yang, Junfeng, Ray, Baishakhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
by: Mathai, Alex, et al.
Published: (2024)
by: Mathai, Alex, et al.
Published: (2024)
CrashFixer: A crash resolution agent for the Linux kernel
by: Mathai, Alex, et al.
Published: (2025)
by: Mathai, Alex, et al.
Published: (2025)
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
by: Gopal, Nikhil, et al.
Published: (2026)
by: Gopal, Nikhil, et al.
Published: (2026)
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025)
by: Xu, Jiakai, et al.
Published: (2025)
SemAgent: A Semantics Aware Program Repair Agent
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
Shabari: Delayed Decision-Making for Faster and Efficient Serverless Functions
by: Sinha, Prasoon, et al.
Published: (2024)
by: Sinha, Prasoon, et al.
Published: (2024)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
by: Pagonas, Nikos, et al.
Published: (2025)
by: Pagonas, Nikos, et al.
Published: (2025)
CRQBench: A Benchmark of Code Reasoning Questions
by: Dinella, Elizabeth, et al.
Published: (2024)
by: Dinella, Elizabeth, et al.
Published: (2024)
BranchBench: Aligning Database Branching with Agentic Demands
by: Ang, Elaine, et al.
Published: (2026)
by: Ang, Elaine, et al.
Published: (2026)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
VineLM: Trie-Based Fine-Grained Control for Agentic Workflows
by: Pagonas, Nikos, et al.
Published: (2026)
by: Pagonas, Nikos, et al.
Published: (2026)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
by: Shi, Sherry, et al.
Published: (2025)
by: Shi, Sherry, et al.
Published: (2025)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
by: Peng, Jinjun, et al.
Published: (2025)
by: Peng, Jinjun, et al.
Published: (2025)
AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent
by: Hua, Wenyue, et al.
Published: (2026)
by: Hua, Wenyue, et al.
Published: (2026)
RubberDuckBench: A Benchmark for AI Coding Assistants
by: Mohammed, Ferida, et al.
Published: (2026)
by: Mohammed, Ferida, et al.
Published: (2026)
Speculative Actions: A Lossless Framework for Faster Agentic Systems
by: Ye, Naimeng, et al.
Published: (2025)
by: Ye, Naimeng, et al.
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
FaultLine: Automated Proof-of-Vulnerability Generation Using LLM Agents
by: Nitin, Vikram, et al.
Published: (2025)
by: Nitin, Vikram, et al.
Published: (2025)
[JGR] - Current waveforms and high-speed camera videos
by: Vukovic, Franjo
Published: (2025)
by: Vukovic, Franjo
Published: (2025)
Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Comment on Revisiting Neural Program Smoothing for Fuzzing
by: She, Dongdong, et al.
Published: (2024)
by: She, Dongdong, et al.
Published: (2024)
Grammatische Terminologie im Kontrast. Einige Überlegungen aus der Sicht des DaF-Unterrichts in Italien
by: Barbara Ivančić
Published: (2010)
by: Barbara Ivančić
Published: (2010)
Case Study: When Enforcement Outruns Accounting
by: Brown, Mya
Published: (2026)
by: Brown, Mya
Published: (2026)
Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance
by: Holzer, Nikolaus, et al.
Published: (2025)
by: Holzer, Nikolaus, et al.
Published: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
Fast Userspace Networking for the Rest of Us
by: Sanaee, Alireza, et al.
Published: (2025)
by: Sanaee, Alireza, et al.
Published: (2025)
Is there a Future in this Past?: Analyzing 15M's Intricate Relation to the Transición
by: Kornetis, Kostis
Published: (2026)
by: Kornetis, Kostis
Published: (2026)
Cutoff in total variation for the shelf shuffle
by: Ottolini, Andrea, et al.
Published: (2024)
by: Ottolini, Andrea, et al.
Published: (2024)
EditLord: Learning Code Transformation Rules for Code Editing
by: Li, Weichen, et al.
Published: (2025)
by: Li, Weichen, et al.
Published: (2025)
Electroweak gauge invariant Higgs multiplets
by: Maniatis, M.
Published: (2024)
by: Maniatis, M.
Published: (2024)
Quality education in the field of sustainability using a statistical analysis
by: Paraschos Maniatis
Published: (2024)
by: Paraschos Maniatis
Published: (2024)
¿Hay un derecho al turismo?
by: Antonio Maniatis
Published: (2019)
by: Antonio Maniatis
Published: (2019)
Wave: Offloading Resource Management to SmartNIC Cores
by: Humphries, Jack Tigar, et al.
Published: (2024)
by: Humphries, Jack Tigar, et al.
Published: (2024)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
by: Kim, Myeongsoo, et al.
Published: (2025)
by: Kim, Myeongsoo, et al.
Published: (2025)
SpecTra: Enhancing the Code Translation Ability of Language Models by Generating Multi-Modal Specifications
by: Nitin, Vikram, et al.
Published: (2024)
by: Nitin, Vikram, et al.
Published: (2024)
Outrunning Big KATs: Efficient Decision Procedures for Variants of GKAT
by: Zhang, Cheng, et al.
Published: (2026)
by: Zhang, Cheng, et al.
Published: (2026)
When Engineering Outruns Intelligence: Rethinking Instruction-Guided Navigation
by: Aghaei, Matin, et al.
Published: (2025)
by: Aghaei, Matin, et al.
Published: (2025)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Similar Items
-
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
by: Mathai, Alex, et al.
Published: (2024) -
CrashFixer: A crash resolution agent for the Linux kernel
by: Mathai, Alex, et al.
Published: (2025) -
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
by: Gopal, Nikhil, et al.
Published: (2026) -
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025) -
SemAgent: A Semantics Aware Program Repair Agent
by: Pabba, Anvith, et al.
Published: (2025)