Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Puigjaner, Albert Gassol, Zacharia, Angelos, Alexis, Kostas
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912868994121728
author Puigjaner, Albert Gassol
Zacharia, Angelos
Alexis, Kostas
author_facet Puigjaner, Albert Gassol
Zacharia, Angelos
Alexis, Kostas
contents Representing and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric reconstructions and can be extended to metric-semantic mapping, they lack a higher level of abstraction and relational reasoning. To address this gap, 3D scene graphs have emerged as a powerful representation for capturing hierarchical structures and object relationships. In this work, we propose an enhanced hierarchical 3D scene graph that integrates open-vocabulary features across multiple abstraction levels and supports object-relational reasoning. Our approach leverages a Vision Language Model (VLM) to infer semantic relationships. Notably, we introduce a task reasoning module that combines Large Language Models (LLM) and a VLM to interpret the scene graph's semantic and relational information, enabling agents to reason about tasks and interact with their environment more intelligently. We validate our method by deploying it on a quadruped robot in multiple environments and tasks, highlighting its ability to reason about them.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02456
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning
Puigjaner, Albert Gassol
Zacharia, Angelos
Alexis, Kostas
Robotics
Representing and understanding 3D environments in a structured manner is crucial for autonomous agents to navigate and reason about their surroundings. While traditional Simultaneous Localization and Mapping (SLAM) methods generate metric reconstructions and can be extended to metric-semantic mapping, they lack a higher level of abstraction and relational reasoning. To address this gap, 3D scene graphs have emerged as a powerful representation for capturing hierarchical structures and object relationships. In this work, we propose an enhanced hierarchical 3D scene graph that integrates open-vocabulary features across multiple abstraction levels and supports object-relational reasoning. Our approach leverages a Vision Language Model (VLM) to infer semantic relationships. Notably, we introduce a task reasoning module that combines Large Language Models (LLM) and a VLM to interpret the scene graph's semantic and relational information, enabling agents to reason about tasks and interact with their environment more intelligently. We validate our method by deploying it on a quadruped robot in multiple environments and tasks, highlighting its ability to reason about them.
title Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning
topic Robotics
url https://arxiv.org/abs/2602.02456