Modeling the Real World with High-Density Visual Particle Dynamics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Whitney, William F., Varley, Jacob, Jain, Deepali, Choromanski, Krzysztof, Singh, Sumeet, Sindhwani, Vikas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917708111544320
author Whitney, William F.
Varley, Jacob
Jain, Deepali
Choromanski, Krzysztof
Singh, Sumeet
Sindhwani, Vikas
author_facet Whitney, William F.
Varley, Jacob
Jain, Deepali
Choromanski, Krzysztof
Singh, Sumeet
Sindhwani, Vikas
contents We present High-Density Visual Particle Dynamics (HD-VPD), a learned world model that can emulate the physical dynamics of real scenes by processing massive latent point clouds containing 100K+ particles. To enable efficiency at this scale, we introduce a novel family of Point Cloud Transformers (PCTs) called Interlacers leveraging intertwined linear-attention Performer layers and graph-based neighbour attention layers. We demonstrate the capabilities of HD-VPD by modeling the dynamics of high degree-of-freedom bi-manual robots with two RGB-D cameras. Compared to the previous graph neural network approach, our Interlacer dynamics is twice as fast with the same prediction quality, and can achieve higher quality using 4x as many particles. We illustrate how HD-VPD can evaluate motion plan quality with robotic box pushing and can grasping tasks. See videos and particle dynamics rendered by HD-VPD at https://sites.google.com/view/hd-vpd.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Modeling the Real World with High-Density Visual Particle Dynamics
Whitney, William F.
Varley, Jacob
Jain, Deepali
Choromanski, Krzysztof
Singh, Sumeet
Sindhwani, Vikas
Machine Learning
Robotics
We present High-Density Visual Particle Dynamics (HD-VPD), a learned world model that can emulate the physical dynamics of real scenes by processing massive latent point clouds containing 100K+ particles. To enable efficiency at this scale, we introduce a novel family of Point Cloud Transformers (PCTs) called Interlacers leveraging intertwined linear-attention Performer layers and graph-based neighbour attention layers. We demonstrate the capabilities of HD-VPD by modeling the dynamics of high degree-of-freedom bi-manual robots with two RGB-D cameras. Compared to the previous graph neural network approach, our Interlacer dynamics is twice as fast with the same prediction quality, and can achieve higher quality using 4x as many particles. We illustrate how HD-VPD can evaluate motion plan quality with robotic box pushing and can grasping tasks. See videos and particle dynamics rendered by HD-VPD at https://sites.google.com/view/hd-vpd.
title Modeling the Real World with High-Density Visual Particle Dynamics
topic Machine Learning
Robotics
url https://arxiv.org/abs/2406.19800