TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Chenghao, Zhang, Jiachen, Li, Chengxuan, Zhou, Zhimu, Wu, Shixin, Huang, Songfang, Duan, Huiling
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909901389824000
author Liu, Chenghao
Zhang, Jiachen
Li, Chengxuan
Zhou, Zhimu
Wu, Shixin
Huang, Songfang
Duan, Huiling
author_facet Liu, Chenghao
Zhang, Jiachen
Li, Chengxuan
Zhou, Zhimu
Wu, Shixin
Huang, Songfang
Duan, Huiling
contents Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This frame-by-frame processing makes models vulnerable to visual noise while ignoring the substantial coherence between consecutive frames in manipulation sequences. We propose Temporal Token Fusion (TTF), a training-free approach that intelligently integrates historical and current visual representations to enhance VLA inference quality. Our method employs dual-dimension detection combining efficient grayscale pixel difference analysis with attention-based semantic relevance assessment, enabling selective temporal token fusion through hard fusion strategies and keyframe anchoring to prevent error accumulation. Comprehensive experiments across LIBERO, SimplerEnv, and real robot tasks demonstrate consistent improvements: 4.0 percentage points average on LIBERO (72.4\% vs 68.4\% baseline), cross-environment validation on SimplerEnv (4.8\% relative improvement), and 8.7\% relative improvement on real robot tasks. Our approach proves model-agnostic, working across OpenVLA and VLA-Cache architectures. Notably, TTF reveals that selective Query matrix reuse in attention mechanisms enhances rather than compromises performance, suggesting promising directions for direct KQV matrix reuse strategies that achieve computational acceleration while improving task success rates.
format Preprint
id arxiv_https___arxiv_org_abs_2508_19257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
Liu, Chenghao
Zhang, Jiachen
Li, Chengxuan
Zhou, Zhimu
Wu, Shixin
Huang, Songfang
Duan, Huiling
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This frame-by-frame processing makes models vulnerable to visual noise while ignoring the substantial coherence between consecutive frames in manipulation sequences. We propose Temporal Token Fusion (TTF), a training-free approach that intelligently integrates historical and current visual representations to enhance VLA inference quality. Our method employs dual-dimension detection combining efficient grayscale pixel difference analysis with attention-based semantic relevance assessment, enabling selective temporal token fusion through hard fusion strategies and keyframe anchoring to prevent error accumulation. Comprehensive experiments across LIBERO, SimplerEnv, and real robot tasks demonstrate consistent improvements: 4.0 percentage points average on LIBERO (72.4\% vs 68.4\% baseline), cross-environment validation on SimplerEnv (4.8\% relative improvement), and 8.7\% relative improvement on real robot tasks. Our approach proves model-agnostic, working across OpenVLA and VLA-Cache architectures. Notably, TTF reveals that selective Query matrix reuse in attention mechanisms enhances rather than compromises performance, suggesting promising directions for direct KQV matrix reuse strategies that achieve computational acceleration while improving task success rates.
title TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2508.19257