OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908738973073408 |
|---|---|
| author | Huang, Huang Liu, Fangchen Fu, Letian Wu, Tingfan Mukadam, Mustafa Malik, Jitendra Goldberg, Ken Abbeel, Pieter |
| author_facet | Huang, Huang Liu, Fangchen Fu, Letian Wu, Tingfan Mukadam, Mustafa Malik, Jitendra Goldberg, Ken Abbeel, Pieter |
| contents | Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visionlanguage models (VLMs) as visual and language features are independently fed into downstream policies, degrading the pre-trained semantic alignments. We propose OTTER, a novel VLA architecture that leverages these existing alignments through explicit, text-aware visual feature extraction. Instead of processing all visual features, OTTER selectively extracts and passes only task-relevant visual features that are semantically aligned with the language instruction to the policy transformer. This allows OTTER to keep the pre-trained vision-language encoders frozen. Thereby, OTTER preserves and utilizes the rich semantic understanding learned from large-scale pre-training, enabling strong zero-shot generalization capabilities. In simulation and real-world experiments, OTTER significantly outperforms existing VLA models, demonstrating strong zeroshot generalization to novel objects and environments. Video, code, checkpoints, and dataset: https://ottervla.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_03734 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction Huang, Huang Liu, Fangchen Fu, Letian Wu, Tingfan Mukadam, Mustafa Malik, Jitendra Goldberg, Ken Abbeel, Pieter Robotics Computer Vision and Pattern Recognition Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visionlanguage models (VLMs) as visual and language features are independently fed into downstream policies, degrading the pre-trained semantic alignments. We propose OTTER, a novel VLA architecture that leverages these existing alignments through explicit, text-aware visual feature extraction. Instead of processing all visual features, OTTER selectively extracts and passes only task-relevant visual features that are semantically aligned with the language instruction to the policy transformer. This allows OTTER to keep the pre-trained vision-language encoders frozen. Thereby, OTTER preserves and utilizes the rich semantic understanding learned from large-scale pre-training, enabling strong zero-shot generalization capabilities. In simulation and real-world experiments, OTTER significantly outperforms existing VLA models, demonstrating strong zeroshot generalization to novel objects and environments. Video, code, checkpoints, and dataset: https://ottervla.github.io/. |
| title | OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.03734 |