Guardado en:
Detalles Bibliográficos
Autores principales: Khanda, Rajat, Baqar, Mohammad, Chakrabarti, Sambuddha, Changdar, Satyasaran
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2507.19555
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912502237888512
author Khanda, Rajat
Baqar, Mohammad
Chakrabarti, Sambuddha
Changdar, Satyasaran
author_facet Khanda, Rajat
Baqar, Mohammad
Chakrabarti, Sambuddha
Changdar, Satyasaran
contents Group Relative Policy Optimization (GRPO) has shown promise in discrete action spaces by eliminating value function dependencies through group-based advantage estimation. However, its application to continuous control remains unexplored, limiting its utility in robotics where continuous actions are essential. This paper presents a theoretical framework extending GRPO to continuous control environments, addressing challenges in high-dimensional action spaces, sparse rewards, and temporal dynamics. Our approach introduces trajectory-based policy clustering, state-aware advantage estimation, and regularized policy updates designed for robotic applications. We provide theoretical analysis of convergence properties and computational complexity, establishing a foundation for future empirical validation in robotic systems including locomotion and manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19555
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Extending Group Relative Policy Optimization to Continuous Control: A Theoretical Framework for Robotic Reinforcement Learning
Khanda, Rajat
Baqar, Mohammad
Chakrabarti, Sambuddha
Changdar, Satyasaran
Robotics
Artificial Intelligence
Group Relative Policy Optimization (GRPO) has shown promise in discrete action spaces by eliminating value function dependencies through group-based advantage estimation. However, its application to continuous control remains unexplored, limiting its utility in robotics where continuous actions are essential. This paper presents a theoretical framework extending GRPO to continuous control environments, addressing challenges in high-dimensional action spaces, sparse rewards, and temporal dynamics. Our approach introduces trajectory-based policy clustering, state-aware advantage estimation, and regularized policy updates designed for robotic applications. We provide theoretical analysis of convergence properties and computational complexity, establishing a foundation for future empirical validation in robotic systems including locomotion and manipulation tasks.
title Extending Group Relative Policy Optimization to Continuous Control: A Theoretical Framework for Robotic Reinforcement Learning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2507.19555