Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Xie, Yiqing, Liu, Emmy, Zhang, Gaokai, Kotalwar, Nachiket, Gandhi, Shubham, Acharya, Sathwik, Wang, Xingyao, Rose, Carolyn, Neubig, Graham, Fried, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
di: Pan, Jiayi, et al.
Pubblicazione: (2024)
di: Pan, Jiayi, et al.
Pubblicazione: (2024)
Training Versatile Coding Agents in Synthetic Environments
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
di: Zhu, Yiqi, et al.
Pubblicazione: (2025)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
di: Xie, Yiqing, et al.
Pubblicazione: (2024)
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
di: Xie, Yiqing, et al.
Pubblicazione: (2025)
di: Xie, Yiqing, et al.
Pubblicazione: (2025)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2024)
di: Wang, Zora Zhiruo, et al.
Pubblicazione: (2024)
MetaLint: Easy-to-Hard Generalization for Code Linting
di: Naik, Atharva, et al.
Pubblicazione: (2025)
di: Naik, Atharva, et al.
Pubblicazione: (2025)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
di: Naik, Atharva, et al.
Pubblicazione: (2024)
di: Naik, Atharva, et al.
Pubblicazione: (2024)
TOM-SWE: User Mental Modeling For Software Engineering Agents
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
di: Zhou, Xuhui, et al.
Pubblicazione: (2025)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
di: Chen, Valerie, et al.
Pubblicazione: (2025)
di: Chen, Valerie, et al.
Pubblicazione: (2025)
Asking What Matters: Reward-Driven Clarification for Software Engineering Tasks
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
di: Vijayvargiya, Sanidhya, et al.
Pubblicazione: (2026)
The BrowserGym Ecosystem for Web Agent Research
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
di: De Chezelles, Thibault Le Sellier, et al.
Pubblicazione: (2024)
CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
di: Kapadnis, Manav Nitin, et al.
Pubblicazione: (2025)
di: Kapadnis, Manav Nitin, et al.
Pubblicazione: (2025)
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
di: Wang, Xingyao, et al.
Pubblicazione: (2025)
di: Wang, Xingyao, et al.
Pubblicazione: (2025)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
V-GameGym: Visual Game Generation for Code Large Language Models
di: Zhang, Wei, et al.
Pubblicazione: (2025)
di: Zhang, Wei, et al.
Pubblicazione: (2025)
ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
di: Gandhi, Shubham, et al.
Pubblicazione: (2025)
Multifaceted Hero Developers and Bug-Fixing Outcomes Across Severity
di: Kumar, Amit, et al.
Pubblicazione: (2026)
di: Kumar, Amit, et al.
Pubblicazione: (2026)
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
Collaborator or Assistant? How AI Coding Agents Partition Work Across Pull Request Lifecycles
di: Jo, Young, et al.
Pubblicazione: (2026)
di: Jo, Young, et al.
Pubblicazione: (2026)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
di: Chi, Wayne, et al.
Pubblicazione: (2025)
di: Chi, Wayne, et al.
Pubblicazione: (2025)
BlueCodeAgent: A Blue Teaming Agent Enabled by Automated Red Teaming for CodeGen AI
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
di: Guo, Chengquan, et al.
Pubblicazione: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
di: Jain, Naman, et al.
Pubblicazione: (2025)
di: Jain, Naman, et al.
Pubblicazione: (2025)
LocAgent: Graph-Guided LLM Agents for Code Localization
di: Chen, Zhaoling, et al.
Pubblicazione: (2025)
di: Chen, Zhaoling, et al.
Pubblicazione: (2025)
Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance
di: Pinna, Giovanni, et al.
Pubblicazione: (2026)
di: Pinna, Giovanni, et al.
Pubblicazione: (2026)
A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
di: Yin, Zhuowen, et al.
Pubblicazione: (2025)
di: Yin, Zhuowen, et al.
Pubblicazione: (2025)
CodeAgent: Autonomous Communicative Agents for Code Review
di: Tang, Xunzhu, et al.
Pubblicazione: (2024)
di: Tang, Xunzhu, et al.
Pubblicazione: (2024)
SolAgent: A Specialized Multi-Agent Framework for Solidity Code Generation
di: Chen, Wei, et al.
Pubblicazione: (2026)
di: Chen, Wei, et al.
Pubblicazione: (2026)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
ReqElicitGym: An Evaluation Environment for Interview Competence in Conversational Requirements Elicitation
di: Jin, Dongming, et al.
Pubblicazione: (2026)
di: Jin, Dongming, et al.
Pubblicazione: (2026)
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
di: Phan, Huy Nhat, et al.
Pubblicazione: (2024)
di: Phan, Huy Nhat, et al.
Pubblicazione: (2024)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
di: Hu, Huimin, et al.
Pubblicazione: (2025)
di: Hu, Huimin, et al.
Pubblicazione: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
di: Ni, Ziyi, et al.
Pubblicazione: (2025)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
di: Chervyakov, Artem, et al.
Pubblicazione: (2025)
di: Chervyakov, Artem, et al.
Pubblicazione: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
di: Zhao, Songwen, et al.
Pubblicazione: (2025)
di: Zhao, Songwen, et al.
Pubblicazione: (2025)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
di: Kovács, Ádám
Pubblicazione: (2026)
di: Kovács, Ádám
Pubblicazione: (2026)
Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks
di: Haque, Md Zahidul, et al.
Pubblicazione: (2026)
di: Haque, Md Zahidul, et al.
Pubblicazione: (2026)
Benchmarking Failures in Tool-Augmented Language Models
di: Treviño, Eduardo, et al.
Pubblicazione: (2025)
di: Treviño, Eduardo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
di: Gandhi, Shubham, et al.
Pubblicazione: (2025) -
Training Software Engineering Agents and Verifiers with SWE-Gym
di: Pan, Jiayi, et al.
Pubblicazione: (2024) -
Training Versatile Coding Agents in Synthetic Environments
di: Zhu, Yiqi, et al.
Pubblicazione: (2025) -
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
di: Xie, Yiqing, et al.
Pubblicazione: (2023) -
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
di: Xie, Yiqing, et al.
Pubblicazione: (2024)