Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
Fuente:
arXiv
Guardado en:
| Autores principales: | Da, Jeff, Wang, Clinton, Deng, Xiang, Ma, Yuntao, Barhate, Nikhil, Hendryx, Sean |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
Learning Goal-Conditioned Representations for Language Reward Models
por: Nath, Vaskar, et al.
Publicado: (2024)
por: Nath, Vaskar, et al.
Publicado: (2024)
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
por: Cheng, Zihao, et al.
Publicado: (2026)
por: Cheng, Zihao, et al.
Publicado: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
Agents in Software Engineering: Survey, Landscape, and Vision
por: Wang, Yanlin, et al.
Publicado: (2024)
por: Wang, Yanlin, et al.
Publicado: (2024)
Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents
por: Kuang, Jiayi, et al.
Publicado: (2025)
por: Kuang, Jiayi, et al.
Publicado: (2025)
Revisiting the Superficial Alignment Hypothesis
por: Raghavendra, Mohit, et al.
Publicado: (2024)
por: Raghavendra, Mohit, et al.
Publicado: (2024)
Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance
por: Ren, Yanwei, et al.
Publicado: (2026)
por: Ren, Yanwei, et al.
Publicado: (2026)
Effective Strategies for Asynchronous Software Engineering Agents
por: Geng, Jiayi, et al.
Publicado: (2026)
por: Geng, Jiayi, et al.
Publicado: (2026)
Detecting RLVR Training Data via Structural Convergence of Reasoning
por: Zhang, Hongbo, et al.
Publicado: (2026)
por: Zhang, Hongbo, et al.
Publicado: (2026)
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback
por: Mehta, Nikhil, et al.
Publicado: (2023)
por: Mehta, Nikhil, et al.
Publicado: (2023)
Agentless: Demystifying LLM-based Software Engineering Agents
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
por: Oh, Jihwan, et al.
Publicado: (2026)
por: Oh, Jihwan, et al.
Publicado: (2026)
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
por: Wang, Clinton J., et al.
Publicado: (2025)
por: Wang, Clinton J., et al.
Publicado: (2025)
MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments
por: Gao, Yicheng, et al.
Publicado: (2026)
por: Gao, Yicheng, et al.
Publicado: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
por: Sharma, Manasi, et al.
Publicado: (2025)
por: Sharma, Manasi, et al.
Publicado: (2025)
Scalable Environments Drive Generalizable Agents
por: Zhang, Jiayi, et al.
Publicado: (2026)
por: Zhang, Jiayi, et al.
Publicado: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
por: Yang, John, et al.
Publicado: (2025)
por: Yang, John, et al.
Publicado: (2025)
Adaptive Milestone Reward for GUI Agents
por: Zheng, Congmin, et al.
Publicado: (2026)
por: Zheng, Congmin, et al.
Publicado: (2026)
Exploring Reasoning Reward Model for Agents
por: Fan, Kaixuan, et al.
Publicado: (2026)
por: Fan, Kaixuan, et al.
Publicado: (2026)
Training Versatile Coding Agents in Synthetic Environments
por: Zhu, Yiqi, et al.
Publicado: (2025)
por: Zhu, Yiqi, et al.
Publicado: (2025)
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
por: Yang, Wang, et al.
Publicado: (2026)
por: Yang, Wang, et al.
Publicado: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
por: Sonwane, Atharv, et al.
Publicado: (2026)
por: Sonwane, Atharv, et al.
Publicado: (2026)
Training Proactive and Personalized LLM Agents
por: Sun, Weiwei, et al.
Publicado: (2025)
por: Sun, Weiwei, et al.
Publicado: (2025)
Training a Generally Curious Agent
por: Tajwar, Fahim, et al.
Publicado: (2025)
por: Tajwar, Fahim, et al.
Publicado: (2025)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
por: Zala, Abhay, et al.
Publicado: (2024)
por: Zala, Abhay, et al.
Publicado: (2024)
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
por: Xue, Xiangyuan, et al.
Publicado: (2025)
por: Xue, Xiangyuan, et al.
Publicado: (2025)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
por: Chen, Peter, et al.
Publicado: (2025)
por: Chen, Peter, et al.
Publicado: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
por: Wang, Peisong, et al.
Publicado: (2025)
por: Wang, Peisong, et al.
Publicado: (2025)
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
por: Guo, Xuehang, et al.
Publicado: (2025)
por: Guo, Xuehang, et al.
Publicado: (2025)
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
por: Men, Tianyi, et al.
Publicado: (2025)
por: Men, Tianyi, et al.
Publicado: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
por: Xia, Yu, et al.
Publicado: (2025)
por: Xia, Yu, et al.
Publicado: (2025)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
por: Yin, Da, et al.
Publicado: (2023)
por: Yin, Da, et al.
Publicado: (2023)
LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble
por: Lee, Yujeong, et al.
Publicado: (2024)
por: Lee, Yujeong, et al.
Publicado: (2024)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
por: Zhu, Yihua, et al.
Publicado: (2026)
por: Zhu, Yihua, et al.
Publicado: (2026)
Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
por: Fang, Tianqing, et al.
Publicado: (2025)
por: Fang, Tianqing, et al.
Publicado: (2025)
ARE: Scaling Up Agent Environments and Evaluations
por: Froger, Romain, et al.
Publicado: (2025)
por: Froger, Romain, et al.
Publicado: (2025)
IntelliAsk: Learning to Ask High-Quality Research Questions via RLVR
por: Sharma, Karun, et al.
Publicado: (2026)
por: Sharma, Karun, et al.
Publicado: (2026)
CRAFT: Grounded Multi-Agent Coordination Under Partial Information
por: Nath, Abhijnan, et al.
Publicado: (2026)
por: Nath, Abhijnan, et al.
Publicado: (2026)
Ejemplares similares
-
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025) -
Learning Goal-Conditioned Representations for Language Reward Models
por: Nath, Vaskar, et al.
Publicado: (2024) -
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
por: Cheng, Zihao, et al.
Publicado: (2026) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025) -
Agents in Software Engineering: Survey, Landscape, and Vision
por: Wang, Yanlin, et al.
Publicado: (2024)