MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Jinyang, Feng, Xiachong, Chen, Qiguang, Zhao, Hanjie, Cheng, Zihui, Bai, Jiesong, Zhou, Jingxuan, Li, Min, Qin, Libo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agent That Debugs: Dynamic State-Guided Vulnerability Repair
von: Liu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Liu, Zhengyao, et al.
Veröffentlicht: (2025)
Guided Debugging of Auto-Translated Code Using Differential Testing
von: Wu, Shengnan, et al.
Veröffentlicht: (2025)
von: Wu, Shengnan, et al.
Veröffentlicht: (2025)
Toward a Better Understanding of Probabilistic Delta Debugging
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2024)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
von: Adnan, Muntasir, et al.
Veröffentlicht: (2025)
von: Adnan, Muntasir, et al.
Veröffentlicht: (2025)
DebugHarness: Emulating Human Dynamic Debugging for Autonomous Program Repair
von: Sun, Maolin, et al.
Veröffentlicht: (2026)
von: Sun, Maolin, et al.
Veröffentlicht: (2026)
Towards Practical and Useful Automated Program Repair for Debugging
von: Xin, Qi, et al.
Veröffentlicht: (2024)
von: Xin, Qi, et al.
Veröffentlicht: (2024)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
DePro: Understanding the Role of LLMs in Debugging Competitive Programming Code
von: Parvez, Nabiha, et al.
Veröffentlicht: (2026)
von: Parvez, Nabiha, et al.
Veröffentlicht: (2026)
Debugging WebAssembly? Put some Whamm on it!
von: Gilbert, Elizabeth, et al.
Veröffentlicht: (2025)
von: Gilbert, Elizabeth, et al.
Veröffentlicht: (2025)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
Revisit Self-Debugging with Self-Generated Tests for Code Generation
von: Chen, Xiancai, et al.
Veröffentlicht: (2025)
von: Chen, Xiancai, et al.
Veröffentlicht: (2025)
Towards Adaptive Software Agents for Debugging
von: Majdoub, Yacine, et al.
Veröffentlicht: (2025)
von: Majdoub, Yacine, et al.
Veröffentlicht: (2025)
Simulated Interactive Debugging
von: Noller, Yannic, et al.
Veröffentlicht: (2025)
von: Noller, Yannic, et al.
Veröffentlicht: (2025)
DebugTA: An LLM-Based Agent for Simplifying Debugging and Teaching in Programming Education
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency
von: Ashrafi, Nazmus, et al.
Veröffentlicht: (2025)
von: Ashrafi, Nazmus, et al.
Veröffentlicht: (2025)
BugSpotter: Automated Generation of Code Debugging Exercises
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
von: Pădurean, Victor-Alexandru, et al.
Veröffentlicht: (2024)
UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
von: Lee, Cheryl, et al.
Veröffentlicht: (2024)
An Exploratory Eye Tracking Study on How Developers Classify and Debug Python Code in Different Paradigms
von: Flint, Samuel W., et al.
Veröffentlicht: (2025)
von: Flint, Samuel W., et al.
Veröffentlicht: (2025)
DebugRepair: Enhancing LLM-Based Automated Program Repair via Self-Directed Debugging
von: Wu, Linhao, et al.
Veröffentlicht: (2026)
von: Wu, Linhao, et al.
Veröffentlicht: (2026)
ISS-Scenario: Scenario-based Testing in CARLA
von: Li, Renjue, et al.
Veröffentlicht: (2024)
von: Li, Renjue, et al.
Veröffentlicht: (2024)
Learning Code-Edit Embedding to Model Student Debugging Behavior
von: Heickal, Hasnain, et al.
Veröffentlicht: (2025)
von: Heickal, Hasnain, et al.
Veröffentlicht: (2025)
Large Language Model Guided Self-Debugging Code Generation
von: Adnan, Muntasir, et al.
Veröffentlicht: (2025)
von: Adnan, Muntasir, et al.
Veröffentlicht: (2025)
ProDebug: An Automated Debugging System for Prolog
von: Brancas, Ricardo, et al.
Veröffentlicht: (2026)
von: Brancas, Ricardo, et al.
Veröffentlicht: (2026)
Timing Analysis Agent: Autonomous Multi-Corner Multi-Mode (MCMM) Timing Debugging with Timing Debug Relation Graph
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
AtPatch: Debugging Transformers via Hot-Fixing Over-Attention
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
Analyzing and Debugging Normative Requirements via Satisfiability Checking
von: Feng, Nick, et al.
Veröffentlicht: (2024)
von: Feng, Nick, et al.
Veröffentlicht: (2024)
WDD: Weighted Delta Debugging
von: Zhou, Xintong, et al.
Veröffentlicht: (2024)
von: Zhou, Xintong, et al.
Veröffentlicht: (2024)
Online and Interactive Bayesian Inference Debugging
von: Nussbaumer, Nathanael, et al.
Veröffentlicht: (2025)
von: Nussbaumer, Nathanael, et al.
Veröffentlicht: (2025)
Learning to Debug: LLM-Organized Knowledge Trees for Solving RTL Assertion Failures
von: Bai, Yunsheng, et al.
Veröffentlicht: (2025)
von: Bai, Yunsheng, et al.
Veröffentlicht: (2025)
Leveraging Print Debugging to Improve Code Generation in Large Language Models
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
von: Hu, Xueyu, et al.
Veröffentlicht: (2024)
Debugging Performance Issues in WebAssembly Runtimes via Mutation-based Inference
von: Zeng, Ruiying, et al.
Veröffentlicht: (2026)
von: Zeng, Ruiying, et al.
Veröffentlicht: (2026)
Industrial Code Quality Benchmarks: Toward Gamification of Software Maintainability
von: Borg, Markus, et al.
Veröffentlicht: (2024)
von: Borg, Markus, et al.
Veröffentlicht: (2024)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2024)
Locating Buggy Segments in Quantum Program Debugging
von: Sato, Naoto, et al.
Veröffentlicht: (2023)
von: Sato, Naoto, et al.
Veröffentlicht: (2023)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
von: Shi, Yuling, et al.
Veröffentlicht: (2024)
Benchmarking ChatGPT, Codeium, and GitHub Copilot: A Comparative Study of AI-Driven Programming and Debugging Assistants
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2024)
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
Structural Verification for Reliable EDA Code Generation without Tool-in-the-Loop Debugging
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2026)
von: Jayasuriya, Dinithi, et al.
Veröffentlicht: (2026)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
von: Manh, Dung Nguyen, et al.
Veröffentlicht: (2024)
von: Manh, Dung Nguyen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Agent That Debugs: Dynamic State-Guided Vulnerability Repair
von: Liu, Zhengyao, et al.
Veröffentlicht: (2025) -
Guided Debugging of Auto-Translated Code Using Differential Testing
von: Wu, Shengnan, et al.
Veröffentlicht: (2025) -
Toward a Better Understanding of Probabilistic Delta Debugging
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2024) -
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026) -
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs
von: Adnan, Muntasir, et al.
Veröffentlicht: (2025)