Vision Language Models Can Parse Floor Plan Maps

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: DeFazio, David, Mehta, Hrudayangam, Wang, Meng, Yang, Ping, Blackburn, Jeremy, Zhang, Shiqi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908672180879360
author DeFazio, David
Mehta, Hrudayangam
Wang, Meng
Yang, Ping
Blackburn, Jeremy
Zhang, Shiqi
author_facet DeFazio, David
Mehta, Hrudayangam
Wang, Meng
Yang, Ping
Blackburn, Jeremy
Zhang, Shiqi
contents Vision language models (VLMs) can simultaneously reason about images and texts to tackle many tasks, from visual question answering to image captioning. This paper focuses on map parsing, a novel task that is unexplored within the VLM context and particularly useful to mobile robots. Map parsing requires understanding not only the labels but also the geometric configurations of a map, i.e., what areas are like and how they are connected. To evaluate the performance of VLMs on map parsing, we prompt VLMs with floor plan maps to generate task plans for complex indoor navigation. Our results demonstrate the remarkable capability of VLMs in map parsing, with a success rate of 0.96 in tasks requiring a sequence of nine navigation actions, e.g., approaching and going through doors. Other than intuitive observations, e.g., VLMs do better in smaller maps and simpler navigation tasks, there was a very interesting observation that its performance drops in large open areas. We provide practical suggestions to address such challenges as validated by our experimental results. Webpage: https://sites.google.com/view/vlm-floorplan/
format Preprint
id arxiv_https___arxiv_org_abs_2409_12842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Vision Language Models Can Parse Floor Plan Maps
DeFazio, David
Mehta, Hrudayangam
Wang, Meng
Yang, Ping
Blackburn, Jeremy
Zhang, Shiqi
Robotics
Artificial Intelligence
Vision language models (VLMs) can simultaneously reason about images and texts to tackle many tasks, from visual question answering to image captioning. This paper focuses on map parsing, a novel task that is unexplored within the VLM context and particularly useful to mobile robots. Map parsing requires understanding not only the labels but also the geometric configurations of a map, i.e., what areas are like and how they are connected. To evaluate the performance of VLMs on map parsing, we prompt VLMs with floor plan maps to generate task plans for complex indoor navigation. Our results demonstrate the remarkable capability of VLMs in map parsing, with a success rate of 0.96 in tasks requiring a sequence of nine navigation actions, e.g., approaching and going through doors. Other than intuitive observations, e.g., VLMs do better in smaller maps and simpler navigation tasks, there was a very interesting observation that its performance drops in large open areas. We provide practical suggestions to address such challenges as validated by our experimental results. Webpage: https://sites.google.com/view/vlm-floorplan/
title Vision Language Models Can Parse Floor Plan Maps
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2409.12842