Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yapağcı, Eray, Öztürk, Yavuz Alp Sencer, Tüzün, Eray |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ImproBR: Bug Report Improver Using LLMs
von: Akyol, Emre Furkan, et al.
Veröffentlicht: (2026)
von: Akyol, Emre Furkan, et al.
Veröffentlicht: (2026)
Past, Present, and Future of Bug Tracking in the Generative AI Era
von: Torun, Utku Boran, et al.
Veröffentlicht: (2025)
von: Torun, Utku Boran, et al.
Veröffentlicht: (2025)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
von: Geruslu, Vehid, et al.
Veröffentlicht: (2026)
von: Geruslu, Vehid, et al.
Veröffentlicht: (2026)
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
von: Gon, Mahmut Furkan, et al.
Veröffentlicht: (2026)
von: Gon, Mahmut Furkan, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models for Code Review
von: Cihan, Umut, et al.
Veröffentlicht: (2025)
von: Cihan, Umut, et al.
Veröffentlicht: (2025)
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review
von: Kamalı, Hüseyin Özgür, et al.
Veröffentlicht: (2026)
von: Kamalı, Hüseyin Özgür, et al.
Veröffentlicht: (2026)
Understanding the Limits of Automated Evaluation for Code Review Bots in Practice
von: Karakaya, Veli, et al.
Veröffentlicht: (2026)
von: Karakaya, Veli, et al.
Veröffentlicht: (2026)
Evaluating the Impact of Data Cleaning on the Quality of Generated Pull Request Descriptions
von: Tire, Kutay, et al.
Veröffentlicht: (2025)
von: Tire, Kutay, et al.
Veröffentlicht: (2025)
Towards Automated Detection of Inline Code Comment Smells
von: Oztas, Ipek, et al.
Veröffentlicht: (2025)
von: Oztas, Ipek, et al.
Veröffentlicht: (2025)
Automated Classification of Human Code Review Comments with Large Language Models
von: Çağlar, Semih, et al.
Veröffentlicht: (2026)
von: Çağlar, Semih, et al.
Veröffentlicht: (2026)
RefExpo: Unveiling Software Project Structures through Advanced Dependency Graph Extraction
von: Haratian, Vahid, et al.
Veröffentlicht: (2024)
von: Haratian, Vahid, et al.
Veröffentlicht: (2024)
PR-Aware Automated Unit Test Generation: Challenges and Opportunities
von: Haratian, Vahid, et al.
Veröffentlicht: (2026)
von: Haratian, Vahid, et al.
Veröffentlicht: (2026)
AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions
von: Torun, Utku Boran, et al.
Veröffentlicht: (2026)
von: Torun, Utku Boran, et al.
Veröffentlicht: (2026)
ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
von: Lu, Pengrui, et al.
Veröffentlicht: (2026)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
von: Liu, Junwei, et al.
Veröffentlicht: (2025)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
von: Cheng, Runxiang, et al.
Veröffentlicht: (2025)
von: Cheng, Runxiang, et al.
Veröffentlicht: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
von: Cheng, Runxiang, et al.
Veröffentlicht: (2026)
von: Cheng, Runxiang, et al.
Veröffentlicht: (2026)
Exploring Large Language Models in Resolving Environment-Related Crash Bugs: Localizing and Repairing
von: Du, Xueying, et al.
Veröffentlicht: (2023)
von: Du, Xueying, et al.
Veröffentlicht: (2023)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
von: Li, Yuanyang, et al.
Veröffentlicht: (2026)
The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
von: Gröpler, Robin, et al.
Veröffentlicht: (2025)
von: Gröpler, Robin, et al.
Veröffentlicht: (2025)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
von: Hu, Ruida, et al.
Veröffentlicht: (2026)
von: Hu, Ruida, et al.
Veröffentlicht: (2026)
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python?
von: Gong, Jianian, et al.
Veröffentlicht: (2024)
von: Gong, Jianian, et al.
Veröffentlicht: (2024)
Argus: Resilience-Oriented Safety Assurance Framework for End-to-End ADSs
von: Wang, Dingji, et al.
Veröffentlicht: (2025)
von: Wang, Dingji, et al.
Veröffentlicht: (2025)
UCAgent: An End-to-End Agent for Block-Level Functional Verification
von: Wang, Junyue, et al.
Veröffentlicht: (2026)
von: Wang, Junyue, et al.
Veröffentlicht: (2026)
Tree-of-Code: A Tree-Structured Exploring Framework for End-to-End Code Generation and Execution in Complex Task Handling
von: Ni, Ziyi, et al.
Veröffentlicht: (2024)
von: Ni, Ziyi, et al.
Veröffentlicht: (2024)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
von: Zheng, Mingwei, et al.
Veröffentlicht: (2025)
von: Zheng, Mingwei, et al.
Veröffentlicht: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
von: Meng, Xiangxin, et al.
Veröffentlicht: (2024)
von: Meng, Xiangxin, et al.
Veröffentlicht: (2024)
MarsCode Agent: AI-native Automated Bug Fixing
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
E2E-REME: Towards End-to-End Microservices Auto-Remediation via Experience-Simulation Reinforcement Fine-Tuning
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
Bus Factor Explorer
von: Klimov, Egor, et al.
Veröffentlicht: (2024)
von: Klimov, Egor, et al.
Veröffentlicht: (2024)
SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models
von: Dikici, Kıvanç Kuzey, et al.
Veröffentlicht: (2026)
von: Dikici, Kıvanç Kuzey, et al.
Veröffentlicht: (2026)
A Serious Game Approach to Introduce the Code Review Practice
von: Baris Ardic, et al.
Veröffentlicht: (2024)
von: Baris Ardic, et al.
Veröffentlicht: (2024)
Bug Analysis Towards Bug Resolution Time Prediction
von: Ozkan, Hasan Yagiz, et al.
Veröffentlicht: (2024)
von: Ozkan, Hasan Yagiz, et al.
Veröffentlicht: (2024)
Energy Consumption of Dataframe Libraries for End-to-End Deep Learning Pipelines:A Comparative Analysis
von: Kumar, Punit, et al.
Veröffentlicht: (2025)
von: Kumar, Punit, et al.
Veröffentlicht: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
von: Guo, JunJia, et al.
Veröffentlicht: (2026)
von: Guo, JunJia, et al.
Veröffentlicht: (2026)
Automated Duplicate Bug Report Detection in Large Open Bug Repositories
von: Laney, Clare E., et al.
Veröffentlicht: (2025)
von: Laney, Clare E., et al.
Veröffentlicht: (2025)
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
von: Wu, Qiushi, et al.
Veröffentlicht: (2025)
von: Wu, Qiushi, et al.
Veröffentlicht: (2025)
Scalable Back-End for an AI-Based Diabetes Prediction Application
von: Radityo, Henry Anand Septian, et al.
Veröffentlicht: (2025)
von: Radityo, Henry Anand Septian, et al.
Veröffentlicht: (2025)
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM Agents
von: Liu, Xiang, et al.
Veröffentlicht: (2026)
von: Liu, Xiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ImproBR: Bug Report Improver Using LLMs
von: Akyol, Emre Furkan, et al.
Veröffentlicht: (2026) -
Past, Present, and Future of Bug Tracking in the Generative AI Era
von: Torun, Utku Boran, et al.
Veröffentlicht: (2025) -
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
von: Geruslu, Vehid, et al.
Veröffentlicht: (2026) -
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
von: Gon, Mahmut Furkan, et al.
Veröffentlicht: (2026) -
Evaluating Large Language Models for Code Review
von: Cihan, Umut, et al.
Veröffentlicht: (2025)