Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Da, Jeff, Wang, Clinton, Deng, Xiang, Ma, Yuntao, Barhate, Nikhil, Hendryx, Sean |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
par: Nath, Vaskar, et autres
Publié: (2025)
par: Nath, Vaskar, et autres
Publié: (2025)
Learning Goal-Conditioned Representations for Language Reward Models
par: Nath, Vaskar, et autres
Publié: (2024)
par: Nath, Vaskar, et autres
Publié: (2024)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
par: Cheng, Zihao, et autres
Publié: (2026)
par: Cheng, Zihao, et autres
Publié: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
par: Gunjal, Anisha, et autres
Publié: (2025)
par: Gunjal, Anisha, et autres
Publié: (2025)
Agents in Software Engineering: Survey, Landscape, and Vision
par: Wang, Yanlin, et autres
Publié: (2024)
par: Wang, Yanlin, et autres
Publié: (2024)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
par: Kuang, Jiayi, et autres
Publié: (2025)
par: Kuang, Jiayi, et autres
Publié: (2025)
Revisiting the Superficial Alignment Hypothesis
par: Raghavendra, Mohit, et autres
Publié: (2024)
par: Raghavendra, Mohit, et autres
Publié: (2024)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
par: Ren, Yanwei, et autres
Publié: (2026)
par: Ren, Yanwei, et autres
Publié: (2026)
Effective Strategies for Asynchronous Software Engineering Agents
par: Geng, Jiayi, et autres
Publié: (2026)
par: Geng, Jiayi, et autres
Publié: (2026)
Detecting RLVR Training Data via Structural Convergence of Reasoning
par: Zhang, Hongbo, et autres
Publié: (2026)
par: Zhang, Hongbo, et autres
Publié: (2026)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
par: Mehta, Nikhil, et autres
Publié: (2023)
par: Mehta, Nikhil, et autres
Publié: (2023)
Agentless: Demystifying LLM-based Software Engineering Agents
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
par: Oh, Jihwan, et autres
Publié: (2026)
par: Oh, Jihwan, et autres
Publié: (2026)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
par: Wang, Clinton J., et autres
Publié: (2025)
par: Wang, Clinton J., et autres
Publié: (2025)
MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments
par: Gao, Yicheng, et autres
Publié: (2026)
par: Gao, Yicheng, et autres
Publié: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
par: Sharma, Manasi, et autres
Publié: (2025)
par: Sharma, Manasi, et autres
Publié: (2025)
Scalable Environments Drive Generalizable Agents
par: Zhang, Jiayi, et autres
Publié: (2026)
par: Zhang, Jiayi, et autres
Publié: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
par: Khalifa, Muhammad, et autres
Publié: (2026)
par: Khalifa, Muhammad, et autres
Publié: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
par: Yang, John, et autres
Publié: (2025)
par: Yang, John, et autres
Publié: (2025)
Adaptive Milestone Reward for GUI Agents
par: Zheng, Congmin, et autres
Publié: (2026)
par: Zheng, Congmin, et autres
Publié: (2026)
Exploring Reasoning Reward Model for Agents
par: Fan, Kaixuan, et autres
Publié: (2026)
par: Fan, Kaixuan, et autres
Publié: (2026)
Training Versatile Coding Agents in Synthetic Environments
par: Zhu, Yiqi, et autres
Publié: (2025)
par: Zhu, Yiqi, et autres
Publié: (2025)
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
par: Yang, Wang, et autres
Publié: (2026)
par: Yang, Wang, et autres
Publié: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
par: Sonwane, Atharv, et autres
Publié: (2026)
par: Sonwane, Atharv, et autres
Publié: (2026)
Training Proactive and Personalized LLM Agents
par: Sun, Weiwei, et autres
Publié: (2025)
par: Sun, Weiwei, et autres
Publié: (2025)
Training a Generally Curious Agent
par: Tajwar, Fahim, et autres
Publié: (2025)
par: Tajwar, Fahim, et autres
Publié: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
par: Zala, Abhay, et autres
Publié: (2024)
par: Zala, Abhay, et autres
Publié: (2024)
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
par: Xue, Xiangyuan, et autres
Publié: (2025)
par: Xue, Xiangyuan, et autres
Publié: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
par: Chen, Peter, et autres
Publié: (2025)
par: Chen, Peter, et autres
Publié: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
par: Wang, Peisong, et autres
Publié: (2025)
par: Wang, Peisong, et autres
Publié: (2025)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
par: Guo, Xuehang, et autres
Publié: (2025)
par: Guo, Xuehang, et autres
Publié: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
par: Men, Tianyi, et autres
Publié: (2025)
par: Men, Tianyi, et autres
Publié: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
par: Xia, Yu, et autres
Publié: (2025)
par: Xia, Yu, et autres
Publié: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
par: Yin, Da, et autres
Publié: (2023)
par: Yin, Da, et autres
Publié: (2023)
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble
par: Lee, Yujeong, et autres
Publié: (2024)
par: Lee, Yujeong, et autres
Publié: (2024)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
par: Zhu, Yihua, et autres
Publié: (2026)
par: Zhu, Yihua, et autres
Publié: (2026)
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
par: Fang, Tianqing, et autres
Publié: (2025)
par: Fang, Tianqing, et autres
Publié: (2025)
ARE: Scaling Up Agent Environments and Evaluations
par: Froger, Romain, et autres
Publié: (2025)
par: Froger, Romain, et autres
Publié: (2025)
IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
par: Sharma, Karun, et autres
Publié: (2026)
par: Sharma, Karun, et autres
Publié: (2026)
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
par: Nath, Abhijnan, et autres
Publié: (2026)
par: Nath, Abhijnan, et autres
Publié: (2026)
Documents similaires
-
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
par: Nath, Vaskar, et autres
Publié: (2025) -
Learning Goal-Conditioned Representations for Language Reward Models
par: Nath, Vaskar, et autres
Publié: (2024) -
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
par: Cheng, Zihao, et autres
Publié: (2026) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
par: Gunjal, Anisha, et autres
Publié: (2025) -
Agents in Software Engineering: Survey, Landscape, and Vision
par: Wang, Yanlin, et autres
Publié: (2024)