PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Huanyu, Wang, Dewei, Wang, Xinmiao, Liu, Xinzhe, Liu, Peng, Bai, Chenjia, Li, Xuelong
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911543604543488
author Li, Huanyu
Wang, Dewei
Wang, Xinmiao
Liu, Xinzhe
Liu, Peng
Bai, Chenjia
Li, Xuelong
author_facet Li, Huanyu
Wang, Dewei
Wang, Xinmiao
Liu, Xinzhe
Liu, Peng
Bai, Chenjia
Li, Xuelong
contents Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and perceptive locomotion, they are constrained by fixed weighting strategies that produce a single suboptimal policy, rather than providing a diverse set of solutions for sophisticated multi-objective control. In this paper, we propose a novel framework leveraging Multi-Objective Reinforcement Learning (MORL) to achieve Preference-Conditioned Humanoid Control (PCHC). Unlike conventional methods that require training a series of policies to approximate the Pareto front, our framework enables a single, preference-conditioned policy to exhibit a wide spectrum of diverse behaviors. To effectively integrate these requirements, we introduce a Beta distribution-based alignment mechanism based on preference vectors modulating a Mixture-of-Experts (MoE) module. We validated our approach on two representative humanoid tasks. Extensive simulations and real-world experiments demonstrate that the proposed framework allows the robot to adaptively shift its objective priorities in real-time based on the input preference condition.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24047
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning
Li, Huanyu
Wang, Dewei
Wang, Xinmiao
Liu, Xinzhe
Liu, Peng
Bai, Chenjia
Li, Xuelong
Robotics
Humanoid robots often need to balance competing objectives, such as maximizing speed while minimizing energy consumption. While current reinforcement learning (RL) methods can master complex skills like fall recovery and perceptive locomotion, they are constrained by fixed weighting strategies that produce a single suboptimal policy, rather than providing a diverse set of solutions for sophisticated multi-objective control. In this paper, we propose a novel framework leveraging Multi-Objective Reinforcement Learning (MORL) to achieve Preference-Conditioned Humanoid Control (PCHC). Unlike conventional methods that require training a series of policies to approximate the Pareto front, our framework enables a single, preference-conditioned policy to exhibit a wide spectrum of diverse behaviors. To effectively integrate these requirements, we introduce a Beta distribution-based alignment mechanism based on preference vectors modulating a Mixture-of-Experts (MoE) module. We validated our approach on two representative humanoid tasks. Extensive simulations and real-world experiments demonstrate that the proposed framework allows the robot to adaptively shift its objective priorities in real-time based on the input preference condition.
title PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning
topic Robotics
url https://arxiv.org/abs/2603.24047