GHIssuemarket: A Sandbox Environment for SWE-Agents Economic Experimentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fouad, Mohamed A., Maia, Marcelo de Almeida |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
par: Yuan, Danlong, et autres
Publié: (2026)
par: Yuan, Danlong, et autres
Publié: (2026)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
par: Sun, Shuang, et autres
Publié: (2026)
par: Sun, Shuang, et autres
Publié: (2026)
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
par: Dihan, Mahir Labib, et autres
Publié: (2026)
par: Dihan, Mahir Labib, et autres
Publié: (2026)
SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
par: Han, Hao, et autres
Publié: (2026)
par: Han, Hao, et autres
Publié: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
par: Prathifkumar, Thanosan, et autres
Publié: (2025)
par: Prathifkumar, Thanosan, et autres
Publié: (2025)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
par: Guan, Hao, et autres
Publié: (2026)
par: Guan, Hao, et autres
Publié: (2026)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
par: Chen, Mouxiang, et autres
Publié: (2026)
par: Chen, Mouxiang, et autres
Publié: (2026)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
par: Ludwig, Nikolai, et autres
Publié: (2026)
par: Ludwig, Nikolai, et autres
Publié: (2026)
U2F: Encouraging SWE-Agent to Seize Novelty without Losing Feasibility
par: Ye, Wencheng, et autres
Publié: (2025)
par: Ye, Wencheng, et autres
Publié: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
par: Pan, Jiayi, et autres
Publié: (2024)
par: Pan, Jiayi, et autres
Publié: (2024)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
par: Gandhi, Shubham, et autres
Publié: (2025)
par: Gandhi, Shubham, et autres
Publié: (2025)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
par: Sahoo, Priyam, et autres
Publié: (2026)
par: Sahoo, Priyam, et autres
Publié: (2026)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
par: Yapağcı, Eray, et autres
Publié: (2025)
par: Yapağcı, Eray, et autres
Publié: (2025)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
par: Xie, Yiqing, et autres
Publié: (2025)
par: Xie, Yiqing, et autres
Publié: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
par: Zhou, Xuhui, et autres
Publié: (2025)
par: Zhou, Xuhui, et autres
Publié: (2025)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
par: Wang, Yuhang, et autres
Publié: (2026)
par: Wang, Yuhang, et autres
Publié: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
par: Raghavendra, Mohit, et autres
Publié: (2026)
par: Raghavendra, Mohit, et autres
Publié: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
Reproduction Test Generation for Java SWE Issues
par: Ahmed, Toufique, et autres
Publié: (2026)
par: Ahmed, Toufique, et autres
Publié: (2026)
SWE-Bench Mobile: Can Large Language Model Agents Develop Industry-Level Mobile Applications?
par: Tian, Muxin, et autres
Publié: (2026)
par: Tian, Muxin, et autres
Publié: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
par: Garg, Spandan, et autres
Publié: (2025)
par: Garg, Spandan, et autres
Publié: (2025)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
par: Liang, Jiarong, et autres
Publié: (2026)
par: Liang, Jiarong, et autres
Publié: (2026)
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
par: Badertdinov, Ibragim, et autres
Publié: (2026)
par: Badertdinov, Ibragim, et autres
Publié: (2026)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
par: Jain, Naman, et autres
Publié: (2025)
par: Jain, Naman, et autres
Publié: (2025)
Playing in the Sandbox: A Study on the Usability of Seccomp
par: Alhindi, Maysara, et autres
Publié: (2025)
par: Alhindi, Maysara, et autres
Publié: (2025)
daVinci-Env: Open SWE Environment Synthesis at Scale
par: Fu, Dayuan, et autres
Publié: (2026)
par: Fu, Dayuan, et autres
Publié: (2026)
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
par: Song, Huatong, et autres
Publié: (2026)
par: Song, Huatong, et autres
Publié: (2026)
SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding
par: Tan, Boyin, et autres
Publié: (2026)
par: Tan, Boyin, et autres
Publié: (2026)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
par: Li, Yuanyang, et autres
Publié: (2026)
par: Li, Yuanyang, et autres
Publié: (2026)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
par: Mhatre, Sanket, et autres
Publié: (2025)
par: Mhatre, Sanket, et autres
Publié: (2025)
Multi-Programming Language Sandbox for LLMs
par: Dou, Shihan, et autres
Publié: (2024)
par: Dou, Shihan, et autres
Publié: (2024)
Sandboxing Adoption in Open Source Ecosystems
par: Alhindi, Maysara, et autres
Publié: (2024)
par: Alhindi, Maysara, et autres
Publié: (2024)
SWE-Bench-CL: Continual Learning for Coding Agents
par: Joshi, Thomas, et autres
Publié: (2025)
par: Joshi, Thomas, et autres
Publié: (2025)
SWE-smith: Scaling Data for Software Engineering Agents
par: Yang, John, et autres
Publié: (2025)
par: Yang, John, et autres
Publié: (2025)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
par: Chen, Zhilong, et autres
Publié: (2025)
par: Chen, Zhilong, et autres
Publié: (2025)
Engineering a Governance-Aware AI Sandbox: Design, Implementation, and Lessons Learned
par: Waseem, Muhammad, et autres
Publié: (2026)
par: Waseem, Muhammad, et autres
Publié: (2026)
Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
par: Ahmed, Toufique, et autres
Publié: (2025)
par: Ahmed, Toufique, et autres
Publié: (2025)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
par: Wang, You, et autres
Publié: (2025)
par: Wang, You, et autres
Publié: (2025)
Can Old Tests Do New Tricks for Resolving SWE Issues?
par: Chen, Yang, et autres
Publié: (2025)
par: Chen, Yang, et autres
Publié: (2025)
What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
par: Martinez, Matias, et autres
Publié: (2026)
par: Martinez, Matias, et autres
Publié: (2026)
Documents similaires
-
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
par: Yuan, Danlong, et autres
Publié: (2026) -
SWE-World: Building Software Engineering Agents in Docker-Free Environments
par: Sun, Shuang, et autres
Publié: (2026) -
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
par: Dihan, Mahir Labib, et autres
Publié: (2026) -
SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
par: Han, Hao, et autres
Publié: (2026) -
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
par: Prathifkumar, Thanosan, et autres
Publié: (2025)