VLM-driven Behavior Tree for Context-aware Task Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wake, Naoki, Kanehira, Atsushi, Takamatsu, Jun, Sasabuchi, Kazuhiro, Ikeuchi, Katsushi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913643512201216
author Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Sasabuchi, Kazuhiro
Ikeuchi, Katsushi
author_facet Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Sasabuchi, Kazuhiro
Ikeuchi, Katsushi
contents The use of Large Language Models (LLMs) for generating Behavior Trees (BTs) has recently gained attention in the robotics community, yet remains in its early stages of development. In this paper, we propose a novel framework that leverages Vision-Language Models (VLMs) to interactively generate and edit BTs that address visual conditions, enabling context-aware robot operations in visually complex environments. A key feature of our approach lies in the conditional control through self-prompted visual conditions. Specifically, the VLM generates BTs with visual condition nodes, where conditions are expressed as free-form text. Another VLM process integrates the text into its prompt and evaluates the conditions against real-world images during robot execution. We validated our framework in a real-world cafe scenario, demonstrating both its feasibility and limitations.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03968
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM-driven Behavior Tree for Context-aware Task Planning
Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Sasabuchi, Kazuhiro
Ikeuchi, Katsushi
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
The use of Large Language Models (LLMs) for generating Behavior Trees (BTs) has recently gained attention in the robotics community, yet remains in its early stages of development. In this paper, we propose a novel framework that leverages Vision-Language Models (VLMs) to interactively generate and edit BTs that address visual conditions, enabling context-aware robot operations in visually complex environments. A key feature of our approach lies in the conditional control through self-prompted visual conditions. Specifically, the VLM generates BTs with visual condition nodes, where conditions are expressed as free-form text. Another VLM process integrates the text into its prompt and evaluates the conditions against real-world images during robot execution. We validated our framework in a real-world cafe scenario, demonstrating both its feasibility and limitations.
title VLM-driven Behavior Tree for Context-aware Task Planning
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2501.03968