CUDABeaver: Benchmarking LLM-Based Automated CUDA Debugging
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Shiyang, Chen, Haoyang, Fazzini, Mattia, Ding, Caiwen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding and Detecting Flaky Builds in GitHub Actions
por: Ge, Wenhao, et al.
Publicado: (2026)
por: Ge, Wenhao, et al.
Publicado: (2026)
Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging
por: Hsieh, Po-Chung, et al.
Publicado: (2025)
por: Hsieh, Po-Chung, et al.
Publicado: (2025)
Code Less to Code More: Streamlining Language Server Protocol and Type System Development for Language Families
por: Bruzzone, Federico, et al.
Publicado: (2025)
por: Bruzzone, Federico, et al.
Publicado: (2025)
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
por: Mughal, Ali Hassaan, et al.
Publicado: (2026)
por: Mughal, Ali Hassaan, et al.
Publicado: (2026)
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)
por: Li, Han, et al.
Publicado: (2026)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
por: Hundal, Rajdeep Singh, et al.
Publicado: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
por: Ramírez, Aurora, et al.
Publicado: (2024)
por: Ramírez, Aurora, et al.
Publicado: (2024)
RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing
por: Guo, Jinyao, et al.
Publicado: (2025)
por: Guo, Jinyao, et al.
Publicado: (2025)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
por: Parris, William M.
Publicado: (2026)
por: Parris, William M.
Publicado: (2026)
DEFault++: Automated Fault Detection, Categorization, and Diagnosis for Transformer Architectures
por: Jahan, Sigma, et al.
Publicado: (2026)
por: Jahan, Sigma, et al.
Publicado: (2026)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
por: Iscan, Mehmet
Publicado: (2026)
por: Iscan, Mehmet
Publicado: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
por: Hu, Yuelin, et al.
Publicado: (2026)
por: Hu, Yuelin, et al.
Publicado: (2026)
L2MAC: Large Language Model Automatic Computer for Extensive Code Generation
por: Holt, Samuel, et al.
Publicado: (2023)
por: Holt, Samuel, et al.
Publicado: (2023)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
por: Jana, Prithwish, et al.
Publicado: (2023)
por: Jana, Prithwish, et al.
Publicado: (2023)
LLMDFA: Analyzing Dataflow in Code with Large Language Models
por: Wang, Chengpeng, et al.
Publicado: (2024)
por: Wang, Chengpeng, et al.
Publicado: (2024)
Code Documentation and Analysis to Secure Software Development
por: Attie, Paul, et al.
Publicado: (2024)
por: Attie, Paul, et al.
Publicado: (2024)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
por: Daneshvar, Seyed Shayan, et al.
Publicado: (2024)
por: Daneshvar, Seyed Shayan, et al.
Publicado: (2024)
DRS-OSS: Practical Diff Risk Scoring with LLMs
por: Sayedsalehi, Ali, et al.
Publicado: (2025)
por: Sayedsalehi, Ali, et al.
Publicado: (2025)
Tests4Py: A Benchmark for System Testing
por: Smytzek, Marius, et al.
Publicado: (2023)
por: Smytzek, Marius, et al.
Publicado: (2023)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
por: Oskooei, Amirkia Rafiei, et al.
Publicado: (2025)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
por: Haseeb, Muhammad
Publicado: (2025)
por: Haseeb, Muhammad
Publicado: (2025)
Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches
por: Wienczkowski, Michael
Publicado: (2026)
por: Wienczkowski, Michael
Publicado: (2026)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
por: Ahmed, Syed Yusuf, et al.
Publicado: (2026)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
por: Wu, Jiaqing, et al.
Publicado: (2026)
por: Wu, Jiaqing, et al.
Publicado: (2026)
Orion: Fuzzing Workflow Automation
por: Bazalii, Max, et al.
Publicado: (2025)
por: Bazalii, Max, et al.
Publicado: (2025)
Highly Interactive Testing for Uninterrupted Development Flow
por: Tropin, Andrew
Publicado: (2025)
por: Tropin, Andrew
Publicado: (2025)
An LSTM-based Test Selection Method for Self-Driving Cars
por: Güllü, Ali, et al.
Publicado: (2025)
por: Güllü, Ali, et al.
Publicado: (2025)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
por: Sartaj, Hassan, et al.
Publicado: (2024)
por: Sartaj, Hassan, et al.
Publicado: (2024)
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
por: Khatchadourian, Raffi, et al.
Publicado: (2025)
por: Khatchadourian, Raffi, et al.
Publicado: (2025)
A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification
por: Odmark, Joshua, et al.
Publicado: (2026)
por: Odmark, Joshua, et al.
Publicado: (2026)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
por: Vargas, Matheus J. T.
Publicado: (2025)
por: Vargas, Matheus J. T.
Publicado: (2025)
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
por: Chen, Jiachi, et al.
Publicado: (2024)
por: Chen, Jiachi, et al.
Publicado: (2024)
Generative AI and the Transformation of Software Development Practices
por: Acharya, Vivek
Publicado: (2025)
por: Acharya, Vivek
Publicado: (2025)
CovRL: Fuzzing JavaScript Engines with Coverage-Guided Reinforcement Learning for LLM-based Mutation
por: Eom, Jueon, et al.
Publicado: (2024)
por: Eom, Jueon, et al.
Publicado: (2024)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
por: Liu, Shunyu, et al.
Publicado: (2025)
por: Liu, Shunyu, et al.
Publicado: (2025)
Spreadsheet Debugging
por: Ayalew, Yirsaw, et al.
Publicado: (2008)
por: Ayalew, Yirsaw, et al.
Publicado: (2008)
Validating Solidity Code Defects using Symbolic and Concrete Execution powered by Large Language Models
por: Susan, Ştefan-Claudiu, et al.
Publicado: (2025)
por: Susan, Ştefan-Claudiu, et al.
Publicado: (2025)
Scattered Forest Search: Smarter Code Space Exploration with LLMs
por: Light, Jonathan, et al.
Publicado: (2024)
por: Light, Jonathan, et al.
Publicado: (2024)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
por: Guo, Dongxin, et al.
Publicado: (2026)
por: Guo, Dongxin, et al.
Publicado: (2026)
Ejemplares similares
-
Understanding and Detecting Flaky Builds in GitHub Actions
por: Ge, Wenhao, et al.
Publicado: (2026) -
Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging
por: Hsieh, Po-Chung, et al.
Publicado: (2025) -
Code Less to Code More: Streamlining Language Server Protocol and Type System Development for Language Families
por: Bruzzone, Federico, et al.
Publicado: (2025) -
Mining Subscenario Refactoring Opportunities in Behaviour-Driven Software Test Suites: ML Classifiers and LLM-Judge Baselines
por: Mughal, Ali Hassaan, et al.
Publicado: (2026) -
CodeTracer: Towards Traceable Agent States
por: Li, Han, et al.
Publicado: (2026)