Learning to Drive Anywhere with Model-Based Reannotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hirose, Noriaki, Ignatova, Lydia, Stachowicz, Kyle, Glossop, Catherine, Levine, Sergey, Shah, Dhruv
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914166863822848
author Hirose, Noriaki
Ignatova, Lydia
Stachowicz, Kyle
Glossop, Catherine
Levine, Sergey
Shah, Dhruv
author_facet Hirose, Noriaki
Ignatova, Lydia
Stachowicz, Kyle
Glossop, Catherine
Levine, Sergey
Shah, Dhruv
contents Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization. To overcome this, we explore leveraging abundant, passively collected data sources, including large volumes of crowd-sourced teleoperation data and unlabeled YouTube videos, despite their potential for lower quality or missing action labels. We propose Model-Based ReAnnotation (MBRA), a framework that utilizes a learned short-horizon, model-based expert model to relabel or generate high-quality actions for these passive datasets. This relabeled data is then distilled into LogoNav, a long-horizon navigation policy conditioned on visual goals or GPS waypoints. We demonstrate that LogoNav, trained using MBRA-processed data, achieves state-of-the-art performance, enabling robust navigation over distances exceeding 300 meters in previously unseen indoor and outdoor environments. Our extensive real-world evaluations, conducted across a fleet of robots (including quadrupeds) in six cities on three continents, validate the policy's ability to generalize and navigate effectively even amidst pedestrians in crowded settings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05592
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Drive Anywhere with Model-Based Reannotation
Hirose, Noriaki
Ignatova, Lydia
Stachowicz, Kyle
Glossop, Catherine
Levine, Sergey
Shah, Dhruv
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Systems and Control
Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization. To overcome this, we explore leveraging abundant, passively collected data sources, including large volumes of crowd-sourced teleoperation data and unlabeled YouTube videos, despite their potential for lower quality or missing action labels. We propose Model-Based ReAnnotation (MBRA), a framework that utilizes a learned short-horizon, model-based expert model to relabel or generate high-quality actions for these passive datasets. This relabeled data is then distilled into LogoNav, a long-horizon navigation policy conditioned on visual goals or GPS waypoints. We demonstrate that LogoNav, trained using MBRA-processed data, achieves state-of-the-art performance, enabling robust navigation over distances exceeding 300 meters in previously unseen indoor and outdoor environments. Our extensive real-world evaluations, conducted across a fleet of robots (including quadrupeds) in six cities on three continents, validate the policy's ability to generalize and navigate effectively even amidst pedestrians in crowded settings.
title Learning to Drive Anywhere with Model-Based Reannotation
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
Systems and Control
url https://arxiv.org/abs/2505.05592