PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Ruiyu, Zhuang, Zheyu, Kragic, Danica, Pokorny, Florian T.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917225221324800
author Wang, Ruiyu
Zhuang, Zheyu
Kragic, Danica
Pokorny, Florian T.
author_facet Wang, Ruiyu
Zhuang, Zheyu
Kragic, Danica
Pokorny, Florian T.
contents Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are typically developed in isolation and often rely on complex pipelines. We introduce PALM (Perception Alignment for Local Manipulation), which leverages the invariance of local action distributions between out-of-distribution (OOD) and demonstrated domains to address these OOD shifts concurrently, without additional input modalities, model changes, or data collection. PALM modularizes the manipulation policy into coarse global components and a local policy for fine-grained actions. We reduce the discrepancy between in-domain and OOD inputs at the local policy level by enforcing local visual focus and consistent proprioceptive representation, allowing the policy to retrieve invariant local actions under OOD conditions. Experiments show that PALM limits OOD performance drops to 8% in simulation and 24% in the real world, compared to 45% and 77% for baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19514
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment
Wang, Ruiyu
Zhuang, Zheyu
Kragic, Danica
Pokorny, Florian T.
Robotics
Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are typically developed in isolation and often rely on complex pipelines. We introduce PALM (Perception Alignment for Local Manipulation), which leverages the invariance of local action distributions between out-of-distribution (OOD) and demonstrated domains to address these OOD shifts concurrently, without additional input modalities, model changes, or data collection. PALM modularizes the manipulation policy into coarse global components and a local policy for fine-grained actions. We reduce the discrepancy between in-domain and OOD inputs at the local policy level by enforcing local visual focus and consistent proprioceptive representation, allowing the policy to retrieve invariant local actions under OOD conditions. Experiments show that PALM limits OOD performance drops to 8% in simulation and 24% in the real world, compared to 45% and 77% for baselines.
title PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment
topic Robotics
url https://arxiv.org/abs/2601.19514