GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Neau, Maëlic, Falomir, Zoe, Santos, Paulo E., Bosser, Anne-Gwenn, Buche, Cédric
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908633434947584
author Neau, Maëlic
Falomir, Zoe
Santos, Paulo E.
Bosser, Anne-Gwenn
Buche, Cédric
author_facet Neau, Maëlic
Falomir, Zoe
Santos, Paulo E.
Bosser, Anne-Gwenn
Buche, Cédric
contents Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
Neau, Maëlic
Falomir, Zoe
Santos, Paulo E.
Bosser, Anne-Gwenn
Buche, Cédric
Robotics
Computer Vision and Pattern Recognition
Deploying autonomous robots that can learn new skills from demonstrations is an important challenge of modern robotics. Existing solutions often apply end-to-end imitation learning with Vision-Language Action (VLA) models or symbolic approaches with Action Model Learning (AML). On the one hand, current VLA models are limited by the lack of high-level symbolic planning, which hinders their abilities in long-horizon tasks. On the other hand, symbolic approaches in AML lack generalization and scalability perspectives. In this paper we present a new neuro-symbolic approach, GraSP-VLA, a framework that uses a Continuous Scene Graph representation to generate a symbolic representation of human demonstrations. This representation is used to generate new planning domains during inference and serves as an orchestrator for low-level VLA policies, scaling up the number of actions that can be reproduced in a row. Our results show that GraSP-VLA is effective for modeling symbolic representations on the task of automatic planning domain generation from observations. In addition, results on real-world experiments show the potential of our Continuous Scene Graph representation to orchestrate low-level VLA policies in long-horizon tasks.
title GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.04357