LACE: Latent Visual Representation for Cross-Embodiment Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jang, Yoo Sung, Ranasinghe, Kanchana, Mata, Cristina, Zhang, Yichi, Mendez-Mendez, Jorge, Ryoo, Michael S.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917502798266368
author Jang, Yoo Sung
Ranasinghe, Kanchana
Mata, Cristina
Zhang, Yichi
Mendez-Mendez, Jorge
Ryoo, Michael S.
author_facet Jang, Yoo Sung
Ranasinghe, Kanchana
Mata, Cristina
Zhang, Yichi
Mendez-Mendez, Jorge
Ryoo, Michael S.
contents Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich inter-class semantics of general objects, we show they fail to establish correspondence between human and robot hands. We propose LACE, a framework that aligns human and robot visual representations in the latent space of these backbones by leveraging correspondences between shared body parts across embodiments as sparse supervision. These annotations can be automatically obtained via forward kinematics, and single robot demonstration is sufficient to train the model. Our semantic alignment loss matches distributions incurred by corresponding features, lifting patch-level supervision to semantic-level alignment, while a Gram loss preserves pretrained feature quality. This alignment enables robot policies to leverage abundant human data when robot demonstrations are scarce: in zero-shot transfer, policies using LACE-DINO outperform those using DINO by a large margin (65\%), with consistent gains in low-data regimes and out-of-distribution environments.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16743
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LACE: Latent Visual Representation for Cross-Embodiment Learning
Jang, Yoo Sung
Ranasinghe, Kanchana
Mata, Cristina
Zhang, Yichi
Mendez-Mendez, Jorge
Ryoo, Michael S.
Robotics
Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich inter-class semantics of general objects, we show they fail to establish correspondence between human and robot hands. We propose LACE, a framework that aligns human and robot visual representations in the latent space of these backbones by leveraging correspondences between shared body parts across embodiments as sparse supervision. These annotations can be automatically obtained via forward kinematics, and single robot demonstration is sufficient to train the model. Our semantic alignment loss matches distributions incurred by corresponding features, lifting patch-level supervision to semantic-level alignment, while a Gram loss preserves pretrained feature quality. This alignment enables robot policies to leverage abundant human data when robot demonstrations are scarce: in zero-shot transfer, policies using LACE-DINO outperform those using DINO by a large margin (65\%), with consistent gains in low-data regimes and out-of-distribution environments.
title LACE: Latent Visual Representation for Cross-Embodiment Learning
topic Robotics
url https://arxiv.org/abs/2605.16743