VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Yixiao, Huo, Mingxiao, Liang, Zhixuan, Du, Yushi, Sun, Lingfeng, Lin, Haotian, Shang, Jinghuan, Peng, Chensheng, Bansal, Mohit, Ding, Mingyu, Tomizuka, Masayoshi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909042650120192
author Wang, Yixiao
Huo, Mingxiao
Liang, Zhixuan
Du, Yushi
Sun, Lingfeng
Lin, Haotian
Shang, Jinghuan
Peng, Chensheng
Bansal, Mohit
Ding, Mingyu
Tomizuka, Masayoshi
author_facet Wang, Yixiao
Huo, Mingxiao
Liang, Zhixuan
Du, Yushi
Sun, Lingfeng
Lin, Haotian
Shang, Jinghuan
Peng, Chensheng
Bansal, Mohit
Ding, Mingyu
Tomizuka, Masayoshi
contents Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only in specific domains, limiting generality across tasks. Distilling multiple VFMs into a unified representation for policy can mitigate this limitation but often yields inflexible task-specific feature selection and requires costly full re-training to incorporate robot-domain knowledge. We propose VER, a Vision Expert transformer for Robot learning. During pretraining, VER distills multiple VFMs into a vision expert library. It then fine-tunes only a lightweight routing network (fewer than 0.4% of parameters) to dynamically select task-relevant experts from the pretrained library for downstream robot tasks. We further introduce Patchwise Expert Routing with Curriculum Top-K Annealing to improve both flexibility and precision of dynamic expert selection. Moreover, VER supports parameter-efficient finetuning for scalable expert utilization and adaptive robot-domain knowledge integration. Across 17 diverse robotic tasks and multiple policy heads, VER achieves state-of-the-art performance. We find that VER reduces large-norm outliers in task-irrelevant regions (e.g., background) and concentrates on task-critical regions. Visualizations and codes can be found in https://yixiaowang7.github.io/ver_page/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
Wang, Yixiao
Huo, Mingxiao
Liang, Zhixuan
Du, Yushi
Sun, Lingfeng
Lin, Haotian
Shang, Jinghuan
Peng, Chensheng
Bansal, Mohit
Ding, Mingyu
Tomizuka, Masayoshi
Robotics
Artificial Intelligence
Machine Learning
Pretrained vision foundation models (VFMs) advance robotic learning via rich visual representations, yet individual VFMs typically excel only in specific domains, limiting generality across tasks. Distilling multiple VFMs into a unified representation for policy can mitigate this limitation but often yields inflexible task-specific feature selection and requires costly full re-training to incorporate robot-domain knowledge. We propose VER, a Vision Expert transformer for Robot learning. During pretraining, VER distills multiple VFMs into a vision expert library. It then fine-tunes only a lightweight routing network (fewer than 0.4% of parameters) to dynamically select task-relevant experts from the pretrained library for downstream robot tasks. We further introduce Patchwise Expert Routing with Curriculum Top-K Annealing to improve both flexibility and precision of dynamic expert selection. Moreover, VER supports parameter-efficient finetuning for scalable expert utilization and adaptive robot-domain knowledge integration. Across 17 diverse robotic tasks and multiple policy heads, VER achieves state-of-the-art performance. We find that VER reduces large-norm outliers in task-irrelevant regions (e.g., background) and concentrates on task-critical regions. Visualizations and codes can be found in https://yixiaowang7.github.io/ver_page/.
title VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.05213