debug-gym: A Text-Based Environment for Interactive Debugging
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Xingdi, Moss, Morgane M, Feghali, Charbel El, Singh, Chinmay, Moldavskaya, Darya, MacPhee, Drew, Caccia, Lucas, Pereira, Matheus, Kim, Minseon, Sordoni, Alessandro, Côté, Marc-Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
by: Sonwane, Atharv, et al.
Published: (2025)
by: Sonwane, Atharv, et al.
Published: (2025)
Gistify! Codebase-Level Understanding via Runtime Execution
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
Learning to Extract Context for Context-Aware LLM Inference
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
TALES: Text Adventure Learning Environment Suite
by: Cui, Christopher Zhang, et al.
Published: (2025)
by: Cui, Christopher Zhang, et al.
Published: (2025)
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
Towards Modular LLMs by Building and Reusing a Library of LoRAs
by: Ostapenko, Oleksiy, et al.
Published: (2024)
by: Ostapenko, Oleksiy, et al.
Published: (2024)
Policy Improvement using Language Feedback Models
by: Zhong, Victor, et al.
Published: (2024)
by: Zhong, Victor, et al.
Published: (2024)
Tracers for debugging and program exploration
by: Chiplunkar, Shardul, et al.
Published: (2026)
by: Chiplunkar, Shardul, et al.
Published: (2026)
Enhancing Agent Learning through World Dynamics Modeling
by: Sun, Zhiyuan, et al.
Published: (2024)
by: Sun, Zhiyuan, et al.
Published: (2024)
Can Language Models Serve as Text-Based World Simulators?
by: Wang, Ruoyao, et al.
Published: (2024)
by: Wang, Ruoyao, et al.
Published: (2024)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
PolyDebug: A Framework for Polyglot Debugging
by: Houdaille, Philémon, et al.
Published: (2025)
by: Houdaille, Philémon, et al.
Published: (2025)
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
NL-Debugging: Exploiting Natural Language as an Intermediate Representation for Code Debugging
by: Zhang, Weiming, et al.
Published: (2025)
by: Zhang, Weiming, et al.
Published: (2025)
WatChat: Explaining perplexing programs by debugging mental models
by: Chandra, Kartik, et al.
Published: (2024)
by: Chandra, Kartik, et al.
Published: (2024)
Language-guided Skill Learning with Temporal Variational Inference
by: Fu, Haotian, et al.
Published: (2024)
by: Fu, Haotian, et al.
Published: (2024)
ProDebug: An Automated Debugging System for Prolog
by: Brancas, Ricardo, et al.
Published: (2026)
by: Brancas, Ricardo, et al.
Published: (2026)
Hear Your Code Fail, Voice-Assisted Debugging for Python
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
by: Amiri, Sayed Mahbub Hasan, et al.
Published: (2025)
Improving Context-Aware Preference Modeling for Language Models
by: Pitis, Silviu, et al.
Published: (2024)
by: Pitis, Silviu, et al.
Published: (2024)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
by: Zhu, Wang Bill, et al.
Published: (2026)
by: Zhu, Wang Bill, et al.
Published: (2026)
Debugging Functional Programs by Interpretation
by: Whitington, John
Published: (2024)
by: Whitington, John
Published: (2024)
Orchard: An Open-Source Agentic Modeling Framework
by: Peng, Baolin, et al.
Published: (2026)
by: Peng, Baolin, et al.
Published: (2026)
Learning to Solve Complex Problems via Dataset Decomposition
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts
by: Kharlamova, Darya, et al.
Published: (2026)
by: Kharlamova, Darya, et al.
Published: (2026)
MdEval: Massively Multilingual Code Debugging
by: Liu, Shukai, et al.
Published: (2024)
by: Liu, Shukai, et al.
Published: (2024)
ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
by: Zhang, Chiyu, et al.
Published: (2025)
by: Zhang, Chiyu, et al.
Published: (2025)
WDD: Weighted Delta Debugging
by: Zhou, Xintong, et al.
Published: (2024)
by: Zhou, Xintong, et al.
Published: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
by: Tian, Runchu, et al.
Published: (2024)
by: Tian, Runchu, et al.
Published: (2024)
Causal-Consistent Reversible Debugging: Improving CauDEr
by: González-Abril, Juan José, et al.
Published: (2024)
by: González-Abril, Juan José, et al.
Published: (2024)
LiveRec: Prototyping Probes by Framing Debug Protocols
by: Döderlein, Jean-Baptiste, et al.
Published: (2024)
by: Döderlein, Jean-Baptiste, et al.
Published: (2024)
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
VinePPO: Refining Credit Assignment in RL Training of LLMs
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
by: Kazemnejad, Amirhossein, et al.
Published: (2024)
Physical Data Embedding for Memory Efficient AI
by: MacPhee, Callen, et al.
Published: (2024)
by: MacPhee, Callen, et al.
Published: (2024)
Exploring hearing loss, cognition and MRI measures of hippocampal substructure in adults
by: Imola X MacPhee, et al.
Published: (2025)
by: Imola X MacPhee, et al.
Published: (2025)
Beyond the Score: Exploring the Associations Between Adverse Childhood Experiences and Electrophysiological Responses to Errors
by: Madeline Fisher, et al.
Published: (2025)
by: Madeline Fisher, et al.
Published: (2025)
Remote Concolic Multiverse Debugging -- Extended Version with Additional Appendices
by: Steevens, Maarten, et al.
Published: (2026)
by: Steevens, Maarten, et al.
Published: (2026)
Similar Items
-
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
by: Sonwane, Atharv, et al.
Published: (2025) -
Gistify! Codebase-Level Understanding via Runtime Execution
by: Lee, Hyunji, et al.
Published: (2025) -
Learning to Extract Context for Context-Aware LLM Inference
by: Kim, Minseon, et al.
Published: (2025) -
Guiding Language Model Reasoning with Planning Tokens
by: Wang, Xinyi, et al.
Published: (2023) -
TALES: Text Adventure Learning Environment Suite
by: Cui, Christopher Zhang, et al.
Published: (2025)