Theia: Distilling Diverse Vision Foundation Models for Robot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Shang, Jinghuan, Schmeckpeper, Karl, May, Brandon B., Minniti, Maria Vittoria, Kelestemur, Tarik, Watkins, David, Herlant, Laura |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sceniris: A Fast Procedural Scene Generation Framework
by: Shang, Jinghuan, et al.
Published: (2025)
by: Shang, Jinghuan, et al.
Published: (2025)
Real-is-Sim: Bridging the Sim-to-Real Gap with a Dynamic Digital Twin
by: Abou-Chakra, Jad, et al.
Published: (2025)
by: Abou-Chakra, Jad, et al.
Published: (2025)
Learning Equivariant Neural-Augmented Object Dynamics From Few Interactions
by: Orozco, Sergio, et al.
Published: (2026)
by: Orozco, Sergio, et al.
Published: (2026)
AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
by: Gong, Ran, et al.
Published: (2025)
by: Gong, Ran, et al.
Published: (2025)
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026)
by: Schroeder, Philip, et al.
Published: (2026)
CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning
by: Li, Xiang, et al.
Published: (2023)
by: Li, Xiang, et al.
Published: (2023)
GenDP: 3D Semantic Fields for Category-Level Generalizable Diffusion Policy
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
On-Robot Reinforcement Learning with Goal-Contrastive Rewards
by: Biza, Ondrej, et al.
Published: (2024)
by: Biza, Ondrej, et al.
Published: (2024)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Reinforcement Learning-Based Monocular Vision Approach for Autonomous UAV Landing
by: Houichime, Tarik, et al.
Published: (2025)
by: Houichime, Tarik, et al.
Published: (2025)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
by: Wang, Yixuan, et al.
Published: (2023)
by: Wang, Yixuan, et al.
Published: (2023)
Vision Foundation Models for Domain Generalisable Cross-View Localisation in Planetary Ground-Aerial Robotic Teams
by: Holden, Lachlan, et al.
Published: (2026)
by: Holden, Lachlan, et al.
Published: (2026)
RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation
by: Chu, Wenhui
Published: (2026)
by: Chu, Wenhui
Published: (2026)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
DeFM: Learning Foundation Representations from Depth for Robotics
by: Patel, Manthan, et al.
Published: (2026)
by: Patel, Manthan, et al.
Published: (2026)
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
by: Wang, Yixiao, et al.
Published: (2025)
by: Wang, Yixiao, et al.
Published: (2025)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
Geometry Meets Vision: Revisiting Pretrained Semantics in Distilled Fields
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
by: Caramia, Donato, et al.
Published: (2025)
by: Caramia, Donato, et al.
Published: (2025)
General Flow as Foundation Affordance for Scalable Robot Learning
by: Yuan, Chengbo, et al.
Published: (2024)
by: Yuan, Chengbo, et al.
Published: (2024)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
The Role of Predictive Uncertainty and Diversity in Embodied AI and Robot Learning
by: Senanayake, Ransalu
Published: (2024)
by: Senanayake, Ransalu
Published: (2024)
Towards an Accurate and Effective Robot Vision (The Problem of Topological Localization for Mobile Robots)
by: Boros, Emanuela
Published: (2025)
by: Boros, Emanuela
Published: (2025)
Mobi-$π$: Mobilizing Your Robot Learning Policy
by: Yang, Jingyun, et al.
Published: (2025)
by: Yang, Jingyun, et al.
Published: (2025)
ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
by: Liao, Yue, et al.
Published: (2025)
by: Liao, Yue, et al.
Published: (2025)
Feasibility of Indoor Frame-Wise Lidar Semantic Segmentation via Distillation from Visual Foundation Model
by: Wu, Haiyang, et al.
Published: (2026)
by: Wu, Haiyang, et al.
Published: (2026)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
by: Yang, Yurou, et al.
Published: (2026)
by: Yang, Yurou, et al.
Published: (2026)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
by: Xing, Youguang, et al.
Published: (2025)
by: Xing, Youguang, et al.
Published: (2025)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
by: Zhao, Ruiteng, et al.
Published: (2026)
by: Zhao, Ruiteng, et al.
Published: (2026)
FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks
by: Zhang, Siqi, et al.
Published: (2025)
by: Zhang, Siqi, et al.
Published: (2025)
Language-guided Robust Navigation for Mobile Robots in Dynamically-changing Environments
by: Simons, Cody, et al.
Published: (2024)
by: Simons, Cody, et al.
Published: (2024)
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception
by: Ma, Jiahao, et al.
Published: (2026)
by: Ma, Jiahao, et al.
Published: (2026)
The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents
by: Yeke, Doguhan, et al.
Published: (2026)
by: Yeke, Doguhan, et al.
Published: (2026)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
by: Guruprasad, Pranav, et al.
Published: (2024)
by: Guruprasad, Pranav, et al.
Published: (2024)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing
by: Chen, Zhiling, et al.
Published: (2026)
by: Chen, Zhiling, et al.
Published: (2026)
Similar Items
-
Sceniris: A Fast Procedural Scene Generation Framework
by: Shang, Jinghuan, et al.
Published: (2025) -
Real-is-Sim: Bridging the Sim-to-Real Gap with a Dynamic Digital Twin
by: Abou-Chakra, Jad, et al.
Published: (2025) -
Learning Equivariant Neural-Augmented Object Dynamics From Few Interactions
by: Orozco, Sergio, et al.
Published: (2026) -
AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
by: Gong, Ran, et al.
Published: (2025) -
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning
by: Schroeder, Philip, et al.
Published: (2026)