Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Changxin, Liang, Junyang, Chang, Yanbin, Xu, Jingzhao, Li, Jianqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916719808741376
author Huang, Changxin
Liang, Junyang
Chang, Yanbin
Xu, Jingzhao
Li, Jianqiang
author_facet Huang, Changxin
Liang, Junyang
Chang, Yanbin
Xu, Jingzhao
Li, Jianqiang
contents Enabling a high-degree-of-freedom robot to learn specific skills is a challenging task due to the complexity of robotic dynamics. Reinforcement learning (RL) has emerged as a promising solution; however, addressing such problems requires the design of multiple reward functions to account for various constraints in robotic motion. Existing approaches typically sum all reward components indiscriminately to optimize the RL value function and policy. We argue that this uniform inclusion of all reward components in policy optimization is inefficient and limits the robot's learning performance. To address this, we propose an Automated Hybrid Reward Scheduling (AHRS) framework based on Large Language Models (LLMs). This paradigm dynamically adjusts the learning intensity of each reward component throughout the policy optimization process, enabling robots to acquire skills in a gradual and structured manner. Specifically, we design a multi-branch value network, where each branch corresponds to a distinct reward component. During policy optimization, each branch is assigned a weight that reflects its importance, and these weights are automatically computed based on rules designed by LLMs. The LLM generates a rule set in advance, derived from the task description, and during training, it selects a weight calculation rule from the library based on language prompts that evaluate the performance of each branch. Experimental results demonstrate that the AHRS method achieves an average 6.48% performance improvement across multiple high-degree-of-freedom robotic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02483
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning
Huang, Changxin
Liang, Junyang
Chang, Yanbin
Xu, Jingzhao
Li, Jianqiang
Robotics
Artificial Intelligence
Enabling a high-degree-of-freedom robot to learn specific skills is a challenging task due to the complexity of robotic dynamics. Reinforcement learning (RL) has emerged as a promising solution; however, addressing such problems requires the design of multiple reward functions to account for various constraints in robotic motion. Existing approaches typically sum all reward components indiscriminately to optimize the RL value function and policy. We argue that this uniform inclusion of all reward components in policy optimization is inefficient and limits the robot's learning performance. To address this, we propose an Automated Hybrid Reward Scheduling (AHRS) framework based on Large Language Models (LLMs). This paradigm dynamically adjusts the learning intensity of each reward component throughout the policy optimization process, enabling robots to acquire skills in a gradual and structured manner. Specifically, we design a multi-branch value network, where each branch corresponds to a distinct reward component. During policy optimization, each branch is assigned a weight that reflects its importance, and these weights are automatically computed based on rules designed by LLMs. The LLM generates a rule set in advance, derived from the task description, and during training, it selects a weight calculation rule from the library based on language prompts that evaluate the performance of each branch. Experimental results demonstrate that the AHRS method achieves an average 6.48% performance improvement across multiple high-degree-of-freedom robotic tasks.
title Automated Hybrid Reward Scheduling via Large Language Models for Robotic Skill Learning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2505.02483