Constrained Group Relative Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Girgis, Roger, de Schaetzen, Rodrigue, Rowe, Luke, Robitaille, Azalée, Pal, Christopher, Paull, Liam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
by: Rowe, Luke, et al.
Published: (2025)
by: Rowe, Luke, et al.
Published: (2025)
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
by: Rowe, Luke, et al.
Published: (2024)
by: Rowe, Luke, et al.
Published: (2024)
Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments
by: Rowe, Luke, et al.
Published: (2025)
by: Rowe, Luke, et al.
Published: (2025)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026)
by: Deng, Jingcheng, et al.
Published: (2026)
State-wise Constrained Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)
by: Zhao, Weiye, et al.
Published: (2023)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
by: Wang, Jialu, et al.
Published: (2026)
by: Wang, Jialu, et al.
Published: (2026)
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
GAGPO: Generalized Advantage Grouped Policy Optimization
by: Zhu, Siyuan, et al.
Published: (2026)
by: Zhu, Siyuan, et al.
Published: (2026)
Reducing Text Bias in Synthetically Generated MCQAs for VLMs in Autonomous Driving
by: Kulgod, Sutej, et al.
Published: (2026)
by: Kulgod, Sutej, et al.
Published: (2026)
LGR2: Language Guided Reward Relabeling for Accelerating Hierarchical Reinforcement Learning
by: Singh, Utsav, et al.
Published: (2024)
by: Singh, Utsav, et al.
Published: (2024)
PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
by: Holk, Simon, et al.
Published: (2024)
by: Holk, Simon, et al.
Published: (2024)
LLMs for Robotic Object Disambiguation
by: Jiang, Connie, et al.
Published: (2024)
by: Jiang, Connie, et al.
Published: (2024)
VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
by: Liu, Xiaoyu, et al.
Published: (2025)
by: Liu, Xiaoyu, et al.
Published: (2025)
Anticipate & Act : Integrating LLMs and Classical Planning for Efficient Task Execution in Household Environments
by: Arora, Raghav, et al.
Published: (2025)
by: Arora, Raghav, et al.
Published: (2025)
Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation Tasks
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
by: Triantafyllidis, Eleftherios, et al.
Published: (2023)
Grounding Large Language Models In Embodied Environment With Imperfect World Models
by: Liu, Haolan, et al.
Published: (2024)
by: Liu, Haolan, et al.
Published: (2024)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
by: Simoni, Marco, et al.
Published: (2025)
by: Simoni, Marco, et al.
Published: (2025)
Constrained Policy Optimization via Sampling-Based Weight-Space Projection
by: Cao, Shengfan, et al.
Published: (2025)
by: Cao, Shengfan, et al.
Published: (2025)
Dynamic Objects Relocalization in Changing Environments with Flow Matching
by: Argenziano, Francesco, et al.
Published: (2025)
by: Argenziano, Francesco, et al.
Published: (2025)
Group Sequence Policy Optimization
by: Zheng, Chujie, et al.
Published: (2025)
by: Zheng, Chujie, et al.
Published: (2025)
Group-Adaptive Threshold Optimization for Robust AI-Generated Text Detection
by: Jung, Minseok, et al.
Published: (2025)
by: Jung, Minseok, et al.
Published: (2025)
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization
by: Deng, Wenlong, et al.
Published: (2025)
by: Deng, Wenlong, et al.
Published: (2025)
Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine
by: Xie, Jiacheng, et al.
Published: (2025)
by: Xie, Jiacheng, et al.
Published: (2025)
Stepwise Alignment for Constrained Language Model Policy Optimization
by: Wachi, Akifumi, et al.
Published: (2024)
by: Wachi, Akifumi, et al.
Published: (2024)
LakotaBERT: A Transformer-based Model for Low Resource Lakota Language
by: Parankusham, Kanishka, et al.
Published: (2025)
by: Parankusham, Kanishka, et al.
Published: (2025)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
Latent Action Pretraining from Videos
by: Ye, Seonghyeon, et al.
Published: (2024)
by: Ye, Seonghyeon, et al.
Published: (2024)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Time-Series JEPA for Predictive Remote Control under Capacity-Limited Networks
by: Girgis, Abanoub M., et al.
Published: (2024)
by: Girgis, Abanoub M., et al.
Published: (2024)
Semantic Communication and Control Co-Design for Multi-Objective Distinct Dynamics
by: Girgis, Abanoub M., et al.
Published: (2024)
by: Girgis, Abanoub M., et al.
Published: (2024)
Diffusion Policy Policy Optimization
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
by: Deng, Mingkai, et al.
Published: (2026)
by: Deng, Mingkai, et al.
Published: (2026)
Constrained Stein Variational Trajectory Optimization
by: Power, Thomas, et al.
Published: (2023)
by: Power, Thomas, et al.
Published: (2023)
Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
by: Huang, Yu Ti
Published: (2025)
by: Huang, Yu Ti
Published: (2025)
Enhancing Speech Instruction Understanding and Disambiguation in Robotics via Speech Prosody
by: Sasu, David, et al.
Published: (2025)
by: Sasu, David, et al.
Published: (2025)
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
by: Chen, Kaiyuan, et al.
Published: (2025)
by: Chen, Kaiyuan, et al.
Published: (2025)
Mental Modeling of Reinforcement Learning Agents by Language Models
by: Lu, Wenhao, et al.
Published: (2024)
by: Lu, Wenhao, et al.
Published: (2024)
Diagnosing Robotics Systems Issues with Large Language Models
by: Herrmann, Jordis Emilia, et al.
Published: (2024)
by: Herrmann, Jordis Emilia, et al.
Published: (2024)
Similar Items
-
Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
by: Rowe, Luke, et al.
Published: (2025) -
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
by: Rowe, Luke, et al.
Published: (2024) -
Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments
by: Rowe, Luke, et al.
Published: (2025) -
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026) -
State-wise Constrained Policy Optimization
by: Zhao, Weiye, et al.
Published: (2023)