Language-Conditioned Offline RL for Multi-Robot Navigation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Morad, Steven, Shankar, Ajay, Blumenkamp, Jan, Prorok, Amanda
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909272828280832
author Morad, Steven
Shankar, Ajay
Blumenkamp, Jan
Prorok, Amanda
author_facet Morad, Steven
Shankar, Ajay
Blumenkamp, Jan
Prorok, Amanda
contents We present a method for developing navigation policies for multi-robot teams that interpret and follow natural language instructions. We condition these policies on embeddings from pretrained Large Language Models (LLMs), and train them via offline reinforcement learning with as little as 20 minutes of randomly-collected data. Experiments on a team of five real robots show that these policies generalize well to unseen commands, indicating an understanding of the LLM latent space. Our method requires no simulators or environment models, and produces low-latency control policies that can be deployed directly to real robots without finetuning. We provide videos of our experiments at https://sites.google.com/view/llm-marl.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language-Conditioned Offline RL for Multi-Robot Navigation
Morad, Steven
Shankar, Ajay
Blumenkamp, Jan
Prorok, Amanda
Robotics
Artificial Intelligence
Machine Learning
We present a method for developing navigation policies for multi-robot teams that interpret and follow natural language instructions. We condition these policies on embeddings from pretrained Large Language Models (LLMs), and train them via offline reinforcement learning with as little as 20 minutes of randomly-collected data. Experiments on a team of five real robots show that these policies generalize well to unseen commands, indicating an understanding of the LLM latent space. Our method requires no simulators or environment models, and produces low-latency control policies that can be deployed directly to real robots without finetuning. We provide videos of our experiments at https://sites.google.com/view/llm-marl.
title Language-Conditioned Offline RL for Multi-Robot Navigation
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.20164