GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908633434947584 |
|---|---|
| author | Neau, Maëlic Falomir, Zoe Santos, Paulo E. Bosser, Anne-Gwenn Buche, Cédric |
| author_facet | Neau, Maëlic Falomir, Zoe Santos, Paulo E. Bosser, Anne-Gwenn Buche, Cédric |
| contents | Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_04357 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies Neau, Maëlic Falomir, Zoe Santos, Paulo E. Bosser, Anne-Gwenn Buche, Cédric Robotics Computer Vision and Pattern Recognition Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks. |
| title | GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.04357 |