Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chahine, Makram, Quach, Alex, Maalouf, Alaa, Wang, Tsun-Hsuan, Rus, Daniela
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908366142439424
author Chahine, Makram
Quach, Alex
Maalouf, Alaa
Wang, Tsun-Hsuan
Rus, Daniela
author_facet Chahine, Makram
Quach, Alex
Maalouf, Alaa
Wang, Tsun-Hsuan
Rus, Daniela
contents End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to generalize beyond their training scenarios, limiting adaptability to new environments, tasks, and concepts. In this work, we investigate the minimal data requirements and architectural adaptations necessary to achieve robust closed-loop performance with vision-based control policies under unseen text instructions and visual distribution shifts. Our findings are synthesized in Flex (Fly lexically), a framework that uses pre-trained Vision Language Models (VLMs) as frozen patch-wise feature extractors, generating spatially aware embeddings that integrate semantic and visual information. We demonstrate the effectiveness of this approach on a quadrotor fly-to-target task, where agents trained via behavior cloning on a small simulated dataset successfully generalize to real-world scenes with diverse novel goals and command formulations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13002
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
Chahine, Makram
Quach, Alex
Maalouf, Alaa
Wang, Tsun-Hsuan
Rus, Daniela
Robotics
Artificial Intelligence
68T40, 68T05, 68T50
I.2.6; I.2.9; I.2.10; I.4.8
End-to-end learning directly maps sensory inputs to actions, creating highly integrated and efficient policies for complex robotics tasks. However, such models often struggle to generalize beyond their training scenarios, limiting adaptability to new environments, tasks, and concepts. In this work, we investigate the minimal data requirements and architectural adaptations necessary to achieve robust closed-loop performance with vision-based control policies under unseen text instructions and visual distribution shifts. Our findings are synthesized in Flex (Fly lexically), a framework that uses pre-trained Vision Language Models (VLMs) as frozen patch-wise feature extractors, generating spatially aware embeddings that integrate semantic and visual information. We demonstrate the effectiveness of this approach on a quadrotor fly-to-target task, where agents trained via behavior cloning on a small simulated dataset successfully generalize to real-world scenes with diverse novel goals and command formulations.
title Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
topic Robotics
Artificial Intelligence
68T40, 68T05, 68T50
I.2.6; I.2.9; I.2.10; I.4.8
url https://arxiv.org/abs/2410.13002