Language-Conditioned Offline RL for Multi-Robot Navigation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909272828280832 |
|---|---|
| author | Morad, Steven Shankar, Ajay Blumenkamp, Jan Prorok, Amanda |
| author_facet | Morad, Steven Shankar, Ajay Blumenkamp, Jan Prorok, Amanda |
| contents | We present a method for developing navigation policies for multi-robot teams that interpret and follow natural language instructions. We condition these policies on embeddings from pretrained Large Language Models (LLMs), and train them via offline reinforcement learning with as little as 20 minutes of randomly-collected data. Experiments on a team of five real robots show that these policies generalize well to unseen commands, indicating an understanding of the LLM latent space. Our method requires no simulators or environment models, and produces low-latency control policies that can be deployed directly to real robots without finetuning. We provide videos of our experiments at https://sites.google.com/view/llm-marl. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_20164 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Language-Conditioned Offline RL for Multi-Robot Navigation Morad, Steven Shankar, Ajay Blumenkamp, Jan Prorok, Amanda Robotics Artificial Intelligence Machine Learning We present a method for developing navigation policies for multi-robot teams that interpret and follow natural language instructions. We condition these policies on embeddings from pretrained Large Language Models (LLMs), and train them via offline reinforcement learning with as little as 20 minutes of randomly-collected data. Experiments on a team of five real robots show that these policies generalize well to unseen commands, indicating an understanding of the LLM latent space. Our method requires no simulators or environment models, and produces low-latency control policies that can be deployed directly to real robots without finetuning. We provide videos of our experiments at https://sites.google.com/view/llm-marl. |
| title | Language-Conditioned Offline RL for Multi-Robot Navigation |
| topic | Robotics Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2407.20164 |