Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Currie, Joel, Migno, Gioele, Piacenti, Enrico, Giannaccini, Maria Elena, Bach, Patric, De Tommaso, Davide, Wykowska, Agnieszka
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909617487872000
author Currie, Joel
Migno, Gioele
Piacenti, Enrico
Giannaccini, Maria Elena
Bach, Patric
De Tommaso, Davide
Wykowska, Agnieszka
author_facet Currie, Joel
Migno, Gioele
Piacenti, Enrico
Giannaccini, Maria Elena
Bach, Patric
De Tommaso, Davide
Wykowska, Agnieszka
contents We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal, we introduce a synthetic dataset, generated in NVIDIA Omniverse, that enables supervised learning for spatial reasoning tasks. Each instance includes an RGB image, a natural language description, and a ground-truth 4X4 transformation matrix representing object pose. We focus on inferring Z-axis distance as a foundational skill, with future extensions targeting full 6 Degrees Of Freedom (DOFs) reasoning. The dataset is publicly available to support further research. This work serves as a foundational step toward embodied AI systems capable of spatial understanding in interactive human-robot scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14366
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
Currie, Joel
Migno, Gioele
Piacenti, Enrico
Giannaccini, Maria Elena
Bach, Patric
De Tommaso, Davide
Wykowska, Agnieszka
Artificial Intelligence
Robotics
We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal, we introduce a synthetic dataset, generated in NVIDIA Omniverse, that enables supervised learning for spatial reasoning tasks. Each instance includes an RGB image, a natural language description, and a ground-truth 4X4 transformation matrix representing object pose. We focus on inferring Z-axis distance as a foundational skill, with future extensions targeting full 6 Degrees Of Freedom (DOFs) reasoning. The dataset is publicly available to support further research. This work serves as a foundational step toward embodied AI systems capable of spatial understanding in interactive human-robot scenarios.
title Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2505.14366