Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
Fuente:
arXiv
Saved in:
| Main Authors: | Yapağcı, Eray, Öztürk, Yavuz Alp Sencer, Tüzün, Eray |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026)
by: Akyol, Emre Furkan, et al.
Published: (2026)
Past, Present, and Future of Bug Tracking in the Generative AI Era
by: Torun, Utku Boran, et al.
Published: (2025)
by: Torun, Utku Boran, et al.
Published: (2025)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
by: Geruslu, Vehid, et al.
Published: (2026)
by: Geruslu, Vehid, et al.
Published: (2026)
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
by: Gon, Mahmut Furkan, et al.
Published: (2026)
by: Gon, Mahmut Furkan, et al.
Published: (2026)
Evaluating Large Language Models for Code Review
by: Cihan, Umut, et al.
Published: (2025)
by: Cihan, Umut, et al.
Published: (2025)
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
Understanding the Limits of Automated Evaluation for Code Review Bots in Practice
by: Karakaya, Veli, et al.
Published: (2026)
by: Karakaya, Veli, et al.
Published: (2026)
Evaluating the Impact of Data Cleaning on the Quality of Generated Pull Request Descriptions
by: Tire, Kutay, et al.
Published: (2025)
by: Tire, Kutay, et al.
Published: (2025)
Towards Automated Detection of Inline Code Comment Smells
by: Oztas, Ipek, et al.
Published: (2025)
by: Oztas, Ipek, et al.
Published: (2025)
Automated Classification of Human Code Review Comments with Large Language Models
by: Çağlar, Semih, et al.
Published: (2026)
by: Çağlar, Semih, et al.
Published: (2026)
RefExpo: Unveiling Software Project Structures through Advanced Dependency Graph Extraction
by: Haratian, Vahid, et al.
Published: (2024)
by: Haratian, Vahid, et al.
Published: (2024)
PR-Aware Automated Unit Test Generation: Challenges and Opportunities
by: Haratian, Vahid, et al.
Published: (2026)
by: Haratian, Vahid, et al.
Published: (2026)
AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions
by: Torun, Utku Boran, et al.
Published: (2026)
by: Torun, Utku Boran, et al.
Published: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
by: Lu, Pengrui, et al.
Published: (2026)
by: Lu, Pengrui, et al.
Published: (2026)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
by: Liu, Junwei, et al.
Published: (2025)
by: Liu, Junwei, et al.
Published: (2025)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
by: Cheng, Runxiang, et al.
Published: (2025)
by: Cheng, Runxiang, et al.
Published: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
Exploring Large Language Models in Resolving Environment-Related Crash Bugs: Localizing and Repairing
by: Du, Xueying, et al.
Published: (2023)
by: Du, Xueying, et al.
Published: (2023)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
by: Li, Yuanyang, et al.
Published: (2026)
by: Li, Yuanyang, et al.
Published: (2026)
The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
by: Gröpler, Robin, et al.
Published: (2025)
by: Gröpler, Robin, et al.
Published: (2025)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
by: Hu, Ruida, et al.
Published: (2026)
by: Hu, Ruida, et al.
Published: (2026)
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python?
by: Gong, Jianian, et al.
Published: (2024)
by: Gong, Jianian, et al.
Published: (2024)
Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
by: Wang, Dingji, et al.
Published: (2025)
by: Wang, Dingji, et al.
Published: (2025)
UCAgent: An End-to-End Agent for Block-Level Functional Verification
by: Wang, Junyue, et al.
Published: (2026)
by: Wang, Junyue, et al.
Published: (2026)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
by: Ni, Ziyi, et al.
Published: (2024)
by: Ni, Ziyi, et al.
Published: (2024)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
by: Meng, Xiangxin, et al.
Published: (2024)
by: Meng, Xiangxin, et al.
Published: (2024)
MarsCode Agent: AI-native Automated Bug Fixing
by: Liu, Yizhou, et al.
Published: (2024)
by: Liu, Yizhou, et al.
Published: (2024)
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Bus Factor Explorer
by: Klimov, Egor, et al.
Published: (2024)
by: Klimov, Egor, et al.
Published: (2024)
SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models
by: Dikici, Kıvanç Kuzey, et al.
Published: (2026)
by: Dikici, Kıvanç Kuzey, et al.
Published: (2026)
A Serious Game Approach to Introduce the Code Review Practice
by: Baris Ardic, et al.
Published: (2024)
by: Baris Ardic, et al.
Published: (2024)
Bug Analysis Towards Bug Resolution Time Prediction
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis
by: Kumar, Punit, et al.
Published: (2025)
by: Kumar, Punit, et al.
Published: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
Automated Duplicate Bug Report Detection in Large Open Bug Repositories
by: Laney, Clare E., et al.
Published: (2025)
by: Laney, Clare E., et al.
Published: (2025)
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
by: Wu, Qiushi, et al.
Published: (2025)
by: Wu, Qiushi, et al.
Published: (2025)
Scalable Back-End for an AI-Based Diabetes Prediction Application
by: Radityo, Henry Anand Septian, et al.
Published: (2025)
by: Radityo, Henry Anand Septian, et al.
Published: (2025)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
by: Liu, Xiang, et al.
Published: (2026)
by: Liu, Xiang, et al.
Published: (2026)
Similar Items
-
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026) -
Past, Present, and Future of Bug Tracking in the Generative AI Era
by: Torun, Utku Boran, et al.
Published: (2025) -
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
by: Geruslu, Vehid, et al.
Published: (2026) -
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
by: Gon, Mahmut Furkan, et al.
Published: (2026) -
Evaluating Large Language Models for Code Review
by: Cihan, Umut, et al.
Published: (2025)