OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Ishika, Goyal, Ankit, Birchfield, Stan, Fox, Dieter, Garg, Animesh, Blukis, Valts
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914162535301120
author Singh, Ishika
Goyal, Ankit
Birchfield, Stan
Fox, Dieter
Garg, Animesh
Blukis, Valts
author_facet Singh, Ishika
Goyal, Ankit
Birchfield, Stan
Fox, Dieter
Garg, Animesh
Blukis, Valts
contents We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural language instructions and one or more RGBD observations to quasi-static robot actions. 3D-aware robot policies achieve state-of-the-art performance on precise robot manipulation tasks, but struggle with generalization to unseen instructions, scenes, and objects. On the other hand, VLAs excel at generalizing across instructions and scenes, but can be sensitive to camera and robot pose variations. We leverage prior knowledge embedded in language and vision foundation models to improve generalization of 3D-aware keyframe policies. OG-VLA unprojects input observations from diverse views into a point cloud which is then rendered from canonical orthographic views, ensuring input view invariance and consistency between input and output spaces. These canonical views are processed with a vision backbone, a Large Language Model (LLM), and an image diffusion model to generate images that encode the next position and orientation of the end-effector on the input scene. Evaluations on the Arnold and Colosseum benchmarks demonstrate state-of-the-art generalization to unseen environments, with over 40% relative improvements while maintaining robust performance in seen settings. We also show real-world adaption in 3 to 5 demonstrations along with strong generalization. Videos and resources at https://og-vla.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2506_01196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
Singh, Ishika
Goyal, Ankit
Birchfield, Stan
Fox, Dieter
Garg, Animesh
Blukis, Valts
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural language instructions and one or more RGBD observations to quasi-static robot actions. 3D-aware robot policies achieve state-of-the-art performance on precise robot manipulation tasks, but struggle with generalization to unseen instructions, scenes, and objects. On the other hand, VLAs excel at generalizing across instructions and scenes, but can be sensitive to camera and robot pose variations. We leverage prior knowledge embedded in language and vision foundation models to improve generalization of 3D-aware keyframe policies. OG-VLA unprojects input observations from diverse views into a point cloud which is then rendered from canonical orthographic views, ensuring input view invariance and consistency between input and output spaces. These canonical views are processed with a vision backbone, a Large Language Model (LLM), and an image diffusion model to generate images that encode the next position and orientation of the end-effector on the input scene. Evaluations on the Arnold and Colosseum benchmarks demonstrate state-of-the-art generalization to unseen environments, with over 40% relative improvements while maintaining robust performance in seen settings. We also show real-world adaption in 3 to 5 demonstrations along with strong generalization. Videos and resources at https://og-vla.github.io/
title OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01196