OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Huang, Liu, Fangchen, Fu, Letian, Wu, Tingfan, Mukadam, Mustafa, Malik, Jitendra, Goldberg, Ken, Abbeel, Pieter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908738973073408
author Huang, Huang
Liu, Fangchen
Fu, Letian
Wu, Tingfan
Mukadam, Mustafa
Malik, Jitendra
Goldberg, Ken
Abbeel, Pieter
author_facet Huang, Huang
Liu, Fangchen
Fu, Letian
Wu, Tingfan
Mukadam, Mustafa
Malik, Jitendra
Goldberg, Ken
Abbeel, Pieter
contents Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visionlanguage models (VLMs) as visual and language features are independently fed into downstream policies, degrading the pre-trained semantic alignments. We propose OTTER, a novel VLA architecture that leverages these existing alignments through explicit, text-aware visual feature extraction. Instead of processing all visual features, OTTER selectively extracts and passes only task-relevant visual features that are semantically aligned with the language instruction to the policy transformer. This allows OTTER to keep the pre-trained vision-language encoders frozen. Thereby, OTTER preserves and utilizes the rich semantic understanding learned from large-scale pre-training, enabling strong zero-shot generalization capabilities. In simulation and real-world experiments, OTTER significantly outperforms existing VLA models, demonstrating strong zeroshot generalization to novel objects and environments. Video, code, checkpoints, and dataset: https://ottervla.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
Huang, Huang
Liu, Fangchen
Fu, Letian
Wu, Tingfan
Mukadam, Mustafa
Malik, Jitendra
Goldberg, Ken
Abbeel, Pieter
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models aim to predict robotic actions based on visual observations and language instructions. Existing approaches require fine-tuning pre-trained visionlanguage models (VLMs) as visual and language features are independently fed into downstream policies, degrading the pre-trained semantic alignments. We propose OTTER, a novel VLA architecture that leverages these existing alignments through explicit, text-aware visual feature extraction. Instead of processing all visual features, OTTER selectively extracts and passes only task-relevant visual features that are semantically aligned with the language instruction to the policy transformer. This allows OTTER to keep the pre-trained vision-language encoders frozen. Thereby, OTTER preserves and utilizes the rich semantic understanding learned from large-scale pre-training, enabling strong zero-shot generalization capabilities. In simulation and real-world experiments, OTTER significantly outperforms existing VLA models, demonstrating strong zeroshot generalization to novel objects and environments. Video, code, checkpoints, and dataset: https://ottervla.github.io/.
title OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.03734