RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Haitian, Zhai, Jixiu, Song, Lei, Bian, Jiang, Liu, Qiang, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
von: Cang, Yueyang, et al.
Veröffentlicht: (2026)
von: Cang, Yueyang, et al.
Veröffentlicht: (2026)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents
von: Waqas, Daud, et al.
Veröffentlicht: (2025)
von: Waqas, Daud, et al.
Veröffentlicht: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
von: Lumer, Elias, et al.
Veröffentlicht: (2025)
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
von: Huang, Han, et al.
Veröffentlicht: (2024)
von: Huang, Han, et al.
Veröffentlicht: (2024)
ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
von: Khandelwal, Dinesh, et al.
Veröffentlicht: (2026)
von: Khandelwal, Dinesh, et al.
Veröffentlicht: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
Optimizing Safe and Aligned Language Generation: A Multi-Objective GRPO Approach
von: Li, Xuying, et al.
Veröffentlicht: (2025)
von: Li, Xuying, et al.
Veröffentlicht: (2025)
A$^2$TGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
von: Chen, Dingwei, et al.
Veröffentlicht: (2026)
von: Chen, Dingwei, et al.
Veröffentlicht: (2026)
Multi-Agent Tool-Integrated Policy Optimization
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
von: Mo, Zhanfeng, et al.
Veröffentlicht: (2025)
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing
von: Zhong, Haitian, et al.
Veröffentlicht: (2025)
von: Zhong, Haitian, et al.
Veröffentlicht: (2025)
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
von: Li, Peiji, et al.
Veröffentlicht: (2026)
von: Li, Peiji, et al.
Veröffentlicht: (2026)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
von: Crouse, Maxwell, et al.
Veröffentlicht: (2026)
von: Crouse, Maxwell, et al.
Veröffentlicht: (2026)
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration
von: Modecrua, Wachiravit, et al.
Veröffentlicht: (2026)
von: Modecrua, Wachiravit, et al.
Veröffentlicht: (2026)
ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
von: Agarwal, Mayank, et al.
Veröffentlicht: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
von: Guan, Shengyue, et al.
Veröffentlicht: (2025)
von: Guan, Shengyue, et al.
Veröffentlicht: (2025)
On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
von: Deng, Wenlong, et al.
Veröffentlicht: (2025)
von: Deng, Wenlong, et al.
Veröffentlicht: (2025)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Xu, Haoyuan, et al.
Veröffentlicht: (2026)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
von: Li, Xu, et al.
Veröffentlicht: (2026)
von: Li, Xu, et al.
Veröffentlicht: (2026)
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
von: Wu, Feijie, et al.
Veröffentlicht: (2025)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
von: Zhu, Gongxi, et al.
Veröffentlicht: (2026)
von: Zhu, Gongxi, et al.
Veröffentlicht: (2026)
Constrained Group Relative Policy Optimization
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
von: Yang, Zixuan, et al.
Veröffentlicht: (2026)
von: Yang, Zixuan, et al.
Veröffentlicht: (2026)
GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
von: Wang, Zhichao
Veröffentlicht: (2025)
von: Wang, Zhichao
Veröffentlicht: (2025)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
Direct Multi-Turn Preference Optimization for Language Agents
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
von: Shi, Wentao, et al.
Veröffentlicht: (2024)
Efficient Tool-Calling Multi-Expert NPC Agent for Commonsense Persona-Grounded Dialogue
von: Nuriyev, Mahammad
Veröffentlicht: (2025)
von: Nuriyev, Mahammad
Veröffentlicht: (2025)
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
von: Wang, Zhebo, et al.
Veröffentlicht: (2026)
von: Wang, Zhebo, et al.
Veröffentlicht: (2026)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
von: Wei, Kangda, et al.
Veröffentlicht: (2026)
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
von: Cang, Yueyang, et al.
Veröffentlicht: (2026) -
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026) -
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025) -
Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents
von: Waqas, Daud, et al.
Veröffentlicht: (2025) -
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)