HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seneviratne, Gershom, An, Jianyu, Ellahy, Sahire, Weerakoon, Kasun, Elnoor, Mohamed Bashir, Kannan, Jonathan Deepak, Sunil, Amogha Thalihalla, Manocha, Dinesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908476956999680
author Seneviratne, Gershom
An, Jianyu
Ellahy, Sahire
Weerakoon, Kasun
Elnoor, Mohamed Bashir
Kannan, Jonathan Deepak
Sunil, Amogha Thalihalla
Manocha, Dinesh
author_facet Seneviratne, Gershom
An, Jianyu
Ellahy, Sahire
Weerakoon, Kasun
Elnoor, Mohamed Bashir
Kannan, Jonathan Deepak
Sunil, Amogha Thalihalla
Manocha, Dinesh
contents In this paper, we introduce HALO, a novel Offline Reward Learning algorithm that quantifies human intuition in navigation into a vision-based reward function for robot navigation. HALO learns a reward model from offline data, leveraging expert trajectories collected from mobile robots. During training, actions are uniformly sampled around a reference action and ranked using preference scores derived from a Boltzmann distribution centered on the preferred action, and shaped based on binary user feedback to intuitive navigation queries. The reward model is trained via the Plackett-Luce loss to align with these ranked preferences. To demonstrate the effectiveness of HALO, we deploy its reward model in two downstream applications: (i) an offline learned policy trained directly on the HALO-derived rewards, and (ii) a model-predictive-control (MPC) based planner that incorporates the HALO reward as an additional cost term. This showcases the versatility of HALO across both learning-based and classical navigation frameworks. Our real-world deployments on a Clearpath Husky across diverse scenarios demonstrate that policies trained with HALO generalize effectively to unseen environments and hardware setups not present in the training data. HALO outperforms state-of-the-art vision-based navigation methods, achieving at least a 33.3% improvement in success rate, a 12.9% reduction in normalized trajectory length, and a 26.6% reduction in Frechet distance compared to human expert trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
Seneviratne, Gershom
An, Jianyu
Ellahy, Sahire
Weerakoon, Kasun
Elnoor, Mohamed Bashir
Kannan, Jonathan Deepak
Sunil, Amogha Thalihalla
Manocha, Dinesh
Robotics
In this paper, we introduce HALO, a novel Offline Reward Learning algorithm that quantifies human intuition in navigation into a vision-based reward function for robot navigation. HALO learns a reward model from offline data, leveraging expert trajectories collected from mobile robots. During training, actions are uniformly sampled around a reference action and ranked using preference scores derived from a Boltzmann distribution centered on the preferred action, and shaped based on binary user feedback to intuitive navigation queries. The reward model is trained via the Plackett-Luce loss to align with these ranked preferences. To demonstrate the effectiveness of HALO, we deploy its reward model in two downstream applications: (i) an offline learned policy trained directly on the HALO-derived rewards, and (ii) a model-predictive-control (MPC) based planner that incorporates the HALO reward as an additional cost term. This showcases the versatility of HALO across both learning-based and classical navigation frameworks. Our real-world deployments on a Clearpath Husky across diverse scenarios demonstrate that policies trained with HALO generalize effectively to unseen environments and hardware setups not present in the training data. HALO outperforms state-of-the-art vision-based navigation methods, achieving at least a 33.3% improvement in success rate, a 12.9% reduction in normalized trajectory length, and a 26.6% reduction in Frechet distance compared to human expert trajectories.
title HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
topic Robotics
url https://arxiv.org/abs/2508.01539