Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nasser, Abdelmoamen, Baba'a, Yousef, Mebrahtu, Murad, Madjid, Nadya Abdel, Dias, Jorge, Khonji, Majid
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913007560294400
author Nasser, Abdelmoamen
Baba'a, Yousef
Mebrahtu, Murad
Madjid, Nadya Abdel
Dias, Jorge
Khonji, Majid
author_facet Nasser, Abdelmoamen
Baba'a, Yousef
Mebrahtu, Murad
Madjid, Nadya Abdel
Dias, Jorge
Khonji, Majid
contents Traditional approaches to off-road autonomy rely on separate models for terrain classification, height estimation, and quantifying slip or slope conditions. Utilizing several models requires training each component separately, having task specific datasets, and fine-tuning. In this work, we present a zero-shot approach leveraging SAM2 for environment segmentation and a vision-language model (VLM) to reason about drivable areas. Our approach involves passing to the VLM both the original image and the segmented image annotated with numeric labels for each mask. The VLM is then prompted to identify which regions, represented by these numeric labels, are drivable. Combined with planning and control modules, this unified framework eliminates the need for explicit terrain-specific models and relies instead on the inherent reasoning capabilities of the VLM. Our approach surpasses state-of-the-art trainable models on high resolution segmentation datasets and enables full stack navigation in our Isaac Sim offroad environment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04564
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs
Nasser, Abdelmoamen
Baba'a, Yousef
Mebrahtu, Murad
Madjid, Nadya Abdel
Dias, Jorge
Khonji, Majid
Robotics
Computer Vision and Pattern Recognition
Traditional approaches to off-road autonomy rely on separate models for terrain classification, height estimation, and quantifying slip or slope conditions. Utilizing several models requires training each component separately, having task specific datasets, and fine-tuning. In this work, we present a zero-shot approach leveraging SAM2 for environment segmentation and a vision-language model (VLM) to reason about drivable areas. Our approach involves passing to the VLM both the original image and the segmented image annotated with numeric labels for each mask. The VLM is then prompted to identify which regions, represented by these numeric labels, are drivable. Combined with planning and control modules, this unified framework eliminates the need for explicit terrain-specific models and relies instead on the inherent reasoning capabilities of the VLM. Our approach surpasses state-of-the-art trainable models on high resolution segmentation datasets and enables full stack navigation in our Isaac Sim offroad environment.
title Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.04564