Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cang, Yueyang, Zhang, Xiaoteng, Zhao, Erlu, Ji, Zehua, Liu, Yuhang, He, Yuchen, Ning, Zhiyuan, Yijun, Chen, Que, Wenge, Shi, Li |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Shared DIFF Transformer
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
DINT Transformer
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
von: Cang, Yueyang, et al.
Veröffentlicht: (2025)
RetCompletion:High-Speed Inference Image Completion with Retentive Network
von: Cang, Yueyang, et al.
Veröffentlicht: (2024)
von: Cang, Yueyang, et al.
Veröffentlicht: (2024)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Intelligent Pathological Diagnosis of Gestational Trophoblastic Diseases via Visual-Language Deep Learning Model
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction
von: Lian, Yongxiang, et al.
Veröffentlicht: (2026)
von: Lian, Yongxiang, et al.
Veröffentlicht: (2026)
GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
M$^{2}$GRPO: Mamba-based Multi-Agent Group Relative Policy Optimization for Biomimetic Underwater Robots Pursuit
von: Feng, Yukai, et al.
Veröffentlicht: (2026)
von: Feng, Yukai, et al.
Veröffentlicht: (2026)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
von: Surana, Rohan, et al.
Veröffentlicht: (2026)
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
von: Han, Kevin, et al.
Veröffentlicht: (2026)
von: Han, Kevin, et al.
Veröffentlicht: (2026)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
von: Kim, Youngeun
Veröffentlicht: (2026)
von: Kim, Youngeun
Veröffentlicht: (2026)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
von: Kwon, Soo Min, et al.
Veröffentlicht: (2026)
von: Kwon, Soo Min, et al.
Veröffentlicht: (2026)
Can KAN Work? Exploring the Potential of Kolmogorov-Arnold Networks in Computer Vision
von: Cang, Yueyang, et al.
Veröffentlicht: (2024)
von: Cang, Yueyang, et al.
Veröffentlicht: (2024)
DaGRPO: Rectifying Gradient Conflict in Reasoning via Distinctiveness-Aware Group Relative Policy Optimization
von: Xie, Xuan, et al.
Veröffentlicht: (2025)
von: Xie, Xuan, et al.
Veröffentlicht: (2025)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
von: Cai, Yuzhu, et al.
Veröffentlicht: (2026)
von: Cai, Yuzhu, et al.
Veröffentlicht: (2026)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
von: Togootogtokh, Enkhtogtokh, et al.
Veröffentlicht: (2025)
Training-Free Group Relative Policy Optimization
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
von: Cai, Yuzheng, et al.
Veröffentlicht: (2025)
The Role of Excitatory Parvalbumin-positive Neurons in the Tectofugal Pathway of Pigeon (Columba livia) Hierarchical Visual Processing
von: Lu, Shan, et al.
Veröffentlicht: (2025)
von: Lu, Shan, et al.
Veröffentlicht: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
GRPO-GCC: Enhancing Cooperation in Spatial Public Goods Games via Group Relative Policy Optimization with Global Cooperation Constraint
von: Yang, Zhaoqilin, et al.
Veröffentlicht: (2025)
von: Yang, Zhaoqilin, et al.
Veröffentlicht: (2025)
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
von: He, Xixiang, et al.
Veröffentlicht: (2026)
von: He, Xixiang, et al.
Veröffentlicht: (2026)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
von: Zhang, Liyu, et al.
Veröffentlicht: (2026)
von: Zhang, Liyu, et al.
Veröffentlicht: (2026)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
von: Zhu, Gongxi, et al.
Veröffentlicht: (2026)
von: Zhu, Gongxi, et al.
Veröffentlicht: (2026)
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
von: He, Dailan, et al.
Veröffentlicht: (2025)
von: He, Dailan, et al.
Veröffentlicht: (2025)
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025)
von: Sane, Soham
Veröffentlicht: (2025)
Constrained Group Relative Policy Optimization
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
von: Girgis, Roger, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Shared DIFF Transformer
von: Cang, Yueyang, et al.
Veröffentlicht: (2025) -
DINT Transformer
von: Cang, Yueyang, et al.
Veröffentlicht: (2025) -
RetCompletion:High-Speed Inference Image Completion with Retentive Network
von: Cang, Yueyang, et al.
Veröffentlicht: (2024) -
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
von: Zhong, Haitian, et al.
Veröffentlicht: (2026) -
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
von: Deng, Jingcheng, et al.
Veröffentlicht: (2026)