Information Filtering via Variational Regularization for Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jinhao, Xia, Wenlong, Wang, Yaojia, Zhou, Zhexuan, Li, Huizhe, Lai, Yichen, Song, Haoming, Gong, Youmin, Mei, Jie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918493302030336
author Zhang, Jinhao
Xia, Wenlong
Wang, Yaojia
Zhou, Zhexuan
Li, Huizhe
Lai, Yichen
Song, Haoming
Gong, Youmin
Mei, Jie
author_facet Zhang, Jinhao
Xia, Wenlong
Wang, Yaojia
Zhou, Zhexuan
Li, Huizhe
Lai, Yichen
Song, Haoming
Gong, Youmin
Mei, Jie
contents Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21926
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Information Filtering via Variational Regularization for Robot Manipulation
Zhang, Jinhao
Xia, Wenlong
Wang, Yaojia
Zhou, Zhexuan
Li, Huizhe
Lai, Yichen
Song, Haoming
Gong, Youmin
Mei, Jie
Robotics
Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments.
title Information Filtering via Variational Regularization for Robot Manipulation
topic Robotics
url https://arxiv.org/abs/2601.21926