Policy Contrastive Decoding for Robotic Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Shihan, Luo, Xu, Zhang, Ji, Xie, Junlin, Song, Jingkuan, Shen, Heng Tao, Gao, Lianli
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918465383694336
author Wu, Shihan
Luo, Xu
Zhang, Ji
Xie, Junlin
Song, Jingkuan
Shen, Heng Tao
Gao, Lianli
author_facet Wu, Shihan
Luo, Xu
Zhang, Ji
Xie, Junlin
Song, Jingkuan
Shen, Heng Tao
Gao, Lianli
contents Robotic foundation models, or generalist robot policies, hold immense potential to enable flexible, general-purpose and dexterous robotic systems. Despite their advancements, our empirical experiments reveal that existing robot policies are prone to learning spurious correlations from pre-training trajectories, adversely affecting their generalization capabilities beyond the training data. To tackle this, we propose a novel Policy Contrastive Decoding (PCD) approach, which redirects the robot policy's focus toward object-relevant visual clues by contrasting action probability distributions derived from original and object-masked visual inputs. As a training-free method, our PCD can be used as a plugin to improve different types of robot policies without needing to finetune or access model weights. We conduct extensive experiments on top of three open-source robot policies, including the autoregressive policy OpenVLA and the diffusion-based policies Octo and $π_0$. The obtained results in both simulation and real-world environments prove PCD's flexibility and effectiveness, e.g., PCD enhances the state-of-the-art policy $π_0$ by 8.9% in the simulation environment and by 108% in the real-world environment. Code and demos are publicly available at: https://koorye.github.io/PCD.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13255
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy Contrastive Decoding for Robotic Foundation Models
Wu, Shihan
Luo, Xu
Zhang, Ji
Xie, Junlin
Song, Jingkuan
Shen, Heng Tao
Gao, Lianli
Robotics
Robotic foundation models, or generalist robot policies, hold immense potential to enable flexible, general-purpose and dexterous robotic systems. Despite their advancements, our empirical experiments reveal that existing robot policies are prone to learning spurious correlations from pre-training trajectories, adversely affecting their generalization capabilities beyond the training data. To tackle this, we propose a novel Policy Contrastive Decoding (PCD) approach, which redirects the robot policy's focus toward object-relevant visual clues by contrasting action probability distributions derived from original and object-masked visual inputs. As a training-free method, our PCD can be used as a plugin to improve different types of robot policies without needing to finetune or access model weights. We conduct extensive experiments on top of three open-source robot policies, including the autoregressive policy OpenVLA and the diffusion-based policies Octo and $π_0$. The obtained results in both simulation and real-world environments prove PCD's flexibility and effectiveness, e.g., PCD enhances the state-of-the-art policy $π_0$ by 8.9% in the simulation environment and by 108% in the real-world environment. Code and demos are publicly available at: https://koorye.github.io/PCD.
title Policy Contrastive Decoding for Robotic Foundation Models
topic Robotics
url https://arxiv.org/abs/2505.13255