Point Bridge: 3D Representations for Cross Domain Policy Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haldar, Siddhant, Johannsmeier, Lars, Pinto, Lerrel, Gupta, Abhishek, Fox, Dieter, Narang, Yashraj, Mandlekar, Ajay
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915890449088512
author Haldar, Siddhant
Johannsmeier, Lars
Pinto, Lerrel
Gupta, Abhishek
Fox, Dieter
Narang, Yashraj
Mandlekar, Ajay
author_facet Haldar, Siddhant
Johannsmeier, Lars
Pinto, Lerrel
Gupta, Abhishek
Fox, Dieter
Narang, Yashraj
Mandlekar, Ajay
contents Robot foundation models are beginning to deliver on the promise of generalist robotic agents, yet progress remains constrained by the scarcity of large-scale real-world manipulation datasets. Simulation and synthetic data generation offer a scalable alternative, but their usefulness is limited by the visual domain gap between simulation and reality. In this work, we present Point Bridge, a framework that leverages unified, domain-agnostic point-based representations to unlock synthetic datasets for zero-shot sim-to-real policy transfer, without explicit visual or object-level alignment. Point Bridge combines automated point-based representation extraction via Vision-Language Models (VLMs), transformer-based policy learning, and efficient inference-time pipelines to train capable real-world manipulation agents using only synthetic data. With additional co-training on small sets of real demonstrations, Point Bridge further improves performance, substantially outperforming prior vision-based sim-and-real co-training methods. It achieves up to 44% gains in zero-shot sim-to-real transfer and up to 66% with limited real data across both single-task and multitask settings. Videos of the robot are best viewed at: https://pointbridge3d.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2601_16212
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Point Bridge: 3D Representations for Cross Domain Policy Learning
Haldar, Siddhant
Johannsmeier, Lars
Pinto, Lerrel
Gupta, Abhishek
Fox, Dieter
Narang, Yashraj
Mandlekar, Ajay
Robotics
Robot foundation models are beginning to deliver on the promise of generalist robotic agents, yet progress remains constrained by the scarcity of large-scale real-world manipulation datasets. Simulation and synthetic data generation offer a scalable alternative, but their usefulness is limited by the visual domain gap between simulation and reality. In this work, we present Point Bridge, a framework that leverages unified, domain-agnostic point-based representations to unlock synthetic datasets for zero-shot sim-to-real policy transfer, without explicit visual or object-level alignment. Point Bridge combines automated point-based representation extraction via Vision-Language Models (VLMs), transformer-based policy learning, and efficient inference-time pipelines to train capable real-world manipulation agents using only synthetic data. With additional co-training on small sets of real demonstrations, Point Bridge further improves performance, substantially outperforming prior vision-based sim-and-real co-training methods. It achieves up to 44% gains in zero-shot sim-to-real transfer and up to 66% with limited real data across both single-task and multitask settings. Videos of the robot are best viewed at: https://pointbridge3d.github.io/
title Point Bridge: 3D Representations for Cross Domain Policy Learning
topic Robotics
url https://arxiv.org/abs/2601.16212