VLM-driven Behavior Tree for Context-aware Task Planning
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913643512201216 |
|---|---|
| author | Wake, Naoki Kanehira, Atsushi Takamatsu, Jun Sasabuchi, Kazuhiro Ikeuchi, Katsushi |
| author_facet | Wake, Naoki Kanehira, Atsushi Takamatsu, Jun Sasabuchi, Kazuhiro Ikeuchi, Katsushi |
| contents | The use of Large Language Models (LLMs) for generating Behavior Trees (BTs) has recently gained attention in the robotics community, yet remains in its early stages of development. In this paper, we propose a novel framework that leverages Vision-Language Models (VLMs) to interactively generate and edit BTs that address visual conditions, enabling context-aware robot operations in visually complex environments. A key feature of our approach lies in the conditional control through self-prompted visual conditions. Specifically, the VLM generates BTs with visual condition nodes, where conditions are expressed as free-form text. Another VLM process integrates the text into its prompt and evaluates the conditions against real-world images during robot execution. We validated our framework in a real-world cafe scenario, demonstrating both its feasibility and limitations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_03968 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VLM-driven Behavior Tree for Context-aware Task Planning Wake, Naoki Kanehira, Atsushi Takamatsu, Jun Sasabuchi, Kazuhiro Ikeuchi, Katsushi Robotics Artificial Intelligence Computer Vision and Pattern Recognition Human-Computer Interaction The use of Large Language Models (LLMs) for generating Behavior Trees (BTs) has recently gained attention in the robotics community, yet remains in its early stages of development. In this paper, we propose a novel framework that leverages Vision-Language Models (VLMs) to interactively generate and edit BTs that address visual conditions, enabling context-aware robot operations in visually complex environments. A key feature of our approach lies in the conditional control through self-prompted visual conditions. Specifically, the VLM generates BTs with visual condition nodes, where conditions are expressed as free-form text. Another VLM process integrates the text into its prompt and evaluates the conditions against real-world images during robot execution. We validated our framework in a real-world cafe scenario, demonstrating both its feasibility and limitations. |
| title | VLM-driven Behavior Tree for Context-aware Task Planning |
| topic | Robotics Artificial Intelligence Computer Vision and Pattern Recognition Human-Computer Interaction |
| url | https://arxiv.org/abs/2501.03968 |