Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Morin, Sacha, Gupta, Kumaraditya, Sandhu, Mahtab, Gauthier, Charlie, Argenziano, Francesco, Ellis, Kirsty, Paull, Liam
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908556454789120
author Morin, Sacha
Gupta, Kumaraditya
Sandhu, Mahtab
Gauthier, Charlie
Argenziano, Francesco
Ellis, Kirsty
Paull, Liam
author_facet Morin, Sacha
Gupta, Kumaraditya
Sandhu, Mahtab
Gauthier, Charlie
Argenziano, Francesco
Ellis, Kirsty
Paull, Liam
contents Executing open-ended natural language queries is a core problem in robotics. While recent advances in imitation learning and vision-language-actions models (VLAs) have enabled promising end-to-end policies, these models struggle when faced with complex instructions and new scenes. An alternative is to design an explicit scene representation as a queryable interface between the robot and the world, using query results to guide downstream motion planning. In this work, we present Agentic Scene Policies (ASP), an agentic framework that leverages the advanced semantic, spatial, and affordance-based querying capabilities of modern scene representations to implement a capable language-conditioned robot policy. ASP can execute open-vocabulary queries in a zero-shot manner by explicitly reasoning about object affordances in the case of more complex skills. Through extensive experiments, we compare ASP with VLAs on tabletop manipulation problems and showcase how ASP can tackle room-level queries through affordance-guided navigation, and a scaled-up scene representation. (Project page: https://montrealrobotics.ca/agentic-scene-policies.github.io/)
format Preprint
id arxiv_https___arxiv_org_abs_2509_19571
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
Morin, Sacha
Gupta, Kumaraditya
Sandhu, Mahtab
Gauthier, Charlie
Argenziano, Francesco
Ellis, Kirsty
Paull, Liam
Robotics
Computer Vision and Pattern Recognition
Executing open-ended natural language queries is a core problem in robotics. While recent advances in imitation learning and vision-language-actions models (VLAs) have enabled promising end-to-end policies, these models struggle when faced with complex instructions and new scenes. An alternative is to design an explicit scene representation as a queryable interface between the robot and the world, using query results to guide downstream motion planning. In this work, we present Agentic Scene Policies (ASP), an agentic framework that leverages the advanced semantic, spatial, and affordance-based querying capabilities of modern scene representations to implement a capable language-conditioned robot policy. ASP can execute open-vocabulary queries in a zero-shot manner by explicitly reasoning about object affordances in the case of more complex skills. Through extensive experiments, we compare ASP with VLAs on tabletop manipulation problems and showcase how ASP can tackle room-level queries through affordance-guided navigation, and a scaled-up scene representation. (Project page: https://montrealrobotics.ca/agentic-scene-policies.github.io/)
title Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.19571