Information Filtering via Variational Regularization for Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918493302030336 |
|---|---|
| author | Zhang, Jinhao Xia, Wenlong Wang, Yaojia Zhou, Zhexuan Li, Huizhe Lai, Yichen Song, Haoming Gong, Youmin Mei, Jie |
| author_facet | Zhang, Jinhao Xia, Wenlong Wang, Yaojia Zhou, Zhexuan Li, Huizhe Lai, Yichen Song, Haoming Gong, Youmin Mei, Jie |
| contents | Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_21926 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Information Filtering via Variational Regularization for Robot Manipulation Zhang, Jinhao Xia, Wenlong Wang, Yaojia Zhou, Zhexuan Li, Huizhe Lai, Yichen Song, Haoming Gong, Youmin Mei, Jie Robotics Diffusion-based visuomotor policies built on 3D visual representations have achieved strong performance in learning complex robotic skills. However, most existing methods employ an oversized denoising decoder. While increasing model capacity can improve denoising, empirical evidence suggests that it also introduces redundancy and noise in intermediate feature blocks. Crucially, we find that randomly masking backbone features in U-Net or skipping intermediate layers in DiT at inference time (without changing training) can improve performance, confirming the presence of task-irrelevant noise in intermediate features. To this end, we propose Variational Regularization (VR), a plug-and-play module that imposes a context-conditioned Gaussian over the noisy features and applies a KL-divergence regularizer, forming an adaptive information bottleneck. Extensive experiments on three simulation benchmarks, RoboTwin2.0, Adroit, and MetaWorld, show that our approach consistently improves task success rates over the baseline for both DP3-UNet and DP3-DiT, achieving new state-of-the-art results. Real-world experiments further demonstrate that our method performs well in practical deployments. |
| title | Information Filtering via Variational Regularization for Robot Manipulation |
| topic | Robotics |
| url | https://arxiv.org/abs/2601.21926 |