OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Huang, Liu, Fangchen, Fu, Letian, Wu, Tingfan, Mukadam, Mustafa, Malik, Jitendra, Goldberg, Ken, Abbeel, Pieter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Touch, Vision, and Language Dataset for Multimodal Alignment
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
PTLD: Sim-to-real Privileged Tactile Latent Distillation for Dexterous Manipulation
by: Chen, Rosy, et al.
Published: (2026)
by: Chen, Rosy, et al.
Published: (2026)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Interactive Task Planning with Language Models
by: Li, Boyi, et al.
Published: (2023)
by: Li, Boyi, et al.
Published: (2023)
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting
by: Liu, Fangchen, et al.
Published: (2024)
by: Liu, Fangchen, et al.
Published: (2024)
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations
by: Yu, Justin, et al.
Published: (2025)
by: Yu, Justin, et al.
Published: (2025)
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
by: Lin, Toru, et al.
Published: (2026)
by: Lin, Toru, et al.
Published: (2026)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
Visual Imitation Enables Contextual Humanoid Control
by: Allshire, Arthur, et al.
Published: (2025)
by: Allshire, Arthur, et al.
Published: (2025)
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
by: Heng, Liang, et al.
Published: (2025)
by: Heng, Liang, et al.
Published: (2025)
DexterityGen: Foundation Controller for Unprecedented Dexterity
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Manipulator as a Tail: Promoting Dynamic Stability for Legged Locomotion
by: Huang, Huang, et al.
Published: (2023)
by: Huang, Huang, et al.
Published: (2023)
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
by: Tang, Yikai, et al.
Published: (2025)
by: Tang, Yikai, et al.
Published: (2025)
Body Transformer: Leveraging Robot Embodiment for Policy Learning
by: Sferrazza, Carmelo, et al.
Published: (2024)
by: Sferrazza, Carmelo, et al.
Published: (2024)
In-Context Imitation Learning via Next-Token Prediction
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
SPIDER: Scalable Physics-Informed Dexterous Retargeting
by: Pan, Chaoyi, et al.
Published: (2025)
by: Pan, Chaoyi, et al.
Published: (2025)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
Cross-Hand Latent Representation for Vision-Language-Action Models
by: Jiang, Guangqi, et al.
Published: (2026)
by: Jiang, Guangqi, et al.
Published: (2026)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
ViTaMIn: Learning Contact-Rich Tasks Through Robot-Free Visuo-Tactile Manipulation Interface
by: Liu, Fangchen, et al.
Published: (2025)
by: Liu, Fangchen, et al.
Published: (2025)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
by: Li, Haosheng, et al.
Published: (2026)
by: Li, Haosheng, et al.
Published: (2026)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
End-to-end RL Improves Dexterous Grasping Policies
by: Singh, Ritvik, et al.
Published: (2025)
by: Singh, Ritvik, et al.
Published: (2025)
TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment
by: Wi, Youngsun, et al.
Published: (2026)
by: Wi, Youngsun, et al.
Published: (2026)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
by: Lou, Haozhe, et al.
Published: (2026)
by: Lou, Haozhe, et al.
Published: (2026)
From LLMs to Actions: Latent Codes as Bridges in Hierarchical Robot Control
by: Shentu, Yide, et al.
Published: (2024)
by: Shentu, Yide, et al.
Published: (2024)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
by: Lin, Haitao, et al.
Published: (2026)
by: Lin, Haitao, et al.
Published: (2026)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
What Matters to You? Towards Visual Representation Alignment for Robot Learning
by: Tian, Ran, et al.
Published: (2023)
by: Tian, Ran, et al.
Published: (2023)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
by: Guo, Wenxuan, et al.
Published: (2026)
by: Guo, Wenxuan, et al.
Published: (2026)
Similar Items
-
A Touch, Vision, and Language Dataset for Multimodal Alignment
by: Fu, Letian, et al.
Published: (2024) -
PTLD: Sim-to-real Privileged Tactile Latent Distillation for Dexterous Manipulation
by: Chen, Rosy, et al.
Published: (2026) -
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025) -
Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm
by: Yin, Zhao-Heng, et al.
Published: (2025) -
Interactive Task Planning with Language Models
by: Li, Boyi, et al.
Published: (2023)