GLaD: Geometric Latent Distillation for Vision-Language-Action Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Minghao, Cao, Meng, Tao, Jiachen, Xu, Rongtao, Yan, Yan, Liang, Xiaodan, Laptev, Ivan, Chang, Xiaojun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915666309677056
author Guo, Minghao
Cao, Meng
Tao, Jiachen
Xu, Rongtao
Yan, Yan
Liang, Xiaodan
Laptev, Ivan
Chang, Xiaojun
author_facet Guo, Minghao
Cao, Meng
Tao, Jiachen
Xu, Rongtao
Yan, Yan
Liang, Xiaodan
Laptev, Ivan
Chang, Xiaojun
contents Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA framework that incorporates 3D geometric priors during pretraining through knowledge distillation. Rather than distilling geometric features solely into the vision encoder, we align the LLM's hidden states corresponding to visual tokens with features from a frozen geometry-aware vision transformer (VGGT), ensuring that geometric understanding is deeply integrated into the multimodal representations that drive action prediction. Pretrained on the Bridge dataset with this geometry distillation mechanism, GLaD achieves 94.1% average success rate across four LIBERO task suites, outperforming UniVLA (92.5%) which uses identical pretraining data. These results validate that geometry-aware pretraining enhances spatial reasoning and policy generalization without requiring explicit depth sensors or 3D annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GLaD: Geometric Latent Distillation for Vision-Language-Action Models
Guo, Minghao
Cao, Meng
Tao, Jiachen
Xu, Rongtao
Yan, Yan
Liang, Xiaodan
Laptev, Ivan
Chang, Xiaojun
Robotics
Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA framework that incorporates 3D geometric priors during pretraining through knowledge distillation. Rather than distilling geometric features solely into the vision encoder, we align the LLM's hidden states corresponding to visual tokens with features from a frozen geometry-aware vision transformer (VGGT), ensuring that geometric understanding is deeply integrated into the multimodal representations that drive action prediction. Pretrained on the Bridge dataset with this geometry distillation mechanism, GLaD achieves 94.1% average success rate across four LIBERO task suites, outperforming UniVLA (92.5%) which uses identical pretraining data. These results validate that geometry-aware pretraining enhances spatial reasoning and policy generalization without requiring explicit depth sensors or 3D annotations.
title GLaD: Geometric Latent Distillation for Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2512.09619