Constrained Robotic Navigation on Preferred Terrains Using LLMs and Speech Instruction: Exploiting the Power of Adverbs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lotfi, Faraz, Faraji, Farnoosh, Kakodkar, Nikhil, Manderson, Travis, Meger, David, Dudek, Gregory
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929301664825344
author Lotfi, Faraz
Faraji, Farnoosh
Kakodkar, Nikhil
Manderson, Travis
Meger, David
Dudek, Gregory
author_facet Lotfi, Faraz
Faraji, Farnoosh
Kakodkar, Nikhil
Manderson, Travis
Meger, David
Dudek, Gregory
contents This paper explores leveraging large language models for map-free off-road navigation using generative AI, reducing the need for traditional data collection and annotation. We propose a method where a robot receives verbal instructions, converted to text through Whisper, and a large language model (LLM) model extracts landmarks, preferred terrains, and crucial adverbs translated into speed settings for constrained navigation. A language-driven semantic segmentation model generates text-based masks for identifying landmarks and terrain types in images. By translating 2D image points to the vehicle's motion plane using camera parameters, an MPC controller can guides the vehicle towards the desired terrain. This approach enhances adaptation to diverse environments and facilitates the use of high-level instructions for navigating complex and challenging terrains.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02294
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Constrained Robotic Navigation on Preferred Terrains Using LLMs and Speech Instruction: Exploiting the Power of Adverbs
Lotfi, Faraz
Faraji, Farnoosh
Kakodkar, Nikhil
Manderson, Travis
Meger, David
Dudek, Gregory
Robotics
Machine Learning
This paper explores leveraging large language models for map-free off-road navigation using generative AI, reducing the need for traditional data collection and annotation. We propose a method where a robot receives verbal instructions, converted to text through Whisper, and a large language model (LLM) model extracts landmarks, preferred terrains, and crucial adverbs translated into speed settings for constrained navigation. A language-driven semantic segmentation model generates text-based masks for identifying landmarks and terrain types in images. By translating 2D image points to the vehicle's motion plane using camera parameters, an MPC controller can guides the vehicle towards the desired terrain. This approach enhances adaptation to diverse environments and facilitates the use of high-level instructions for navigating complex and challenging terrains.
title Constrained Robotic Navigation on Preferred Terrains Using LLMs and Speech Instruction: Exploiting the Power of Adverbs
topic Robotics
Machine Learning
url https://arxiv.org/abs/2404.02294