Test-Time Training for Visual Foresight Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Sangwu, Kim, Wonjoong, In, Yeonjun, Kim, Sein, Kang, Hongseok, Park, Chanyoung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911664098508800
author Park, Sangwu
Kim, Wonjoong
In, Yeonjun
Kim, Sein
Kang, Hongseok
Park, Chanyoung
author_facet Park, Sangwu
Kim, Wonjoong
In, Yeonjun
Kim, Sein
Kang, Hongseok
Park, Chanyoung
contents Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inherent design of VF-VLA makes it particularly vulnerable to out-of-distribution (OOD) shifts. Because the quality of action directly depends on the accuracy of the predicted future visual information, OOD conditions affect both stages at once. To address this vulnerability, we propose Test-Time Training Visual Foresight VLA ($T^3$VF), a test-time training approach motivated by the observation that the predicted future image and its subsequent observation form a natural supervision pair. To further address the practical challenges that arise from indiscriminate test-time updates, we introduce an adaptive update filtering mechanism. Empirically, $T^3$VF mitigates the OOD vulnerability of VF-VLA at a modest additional inference cost, without requiring any architectural modification or auxiliary modules.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08215
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Test-Time Training for Visual Foresight Vision-Language-Action Models
Park, Sangwu
Kim, Wonjoong
In, Yeonjun
Kim, Sein
Kang, Hongseok
Park, Chanyoung
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inherent design of VF-VLA makes it particularly vulnerable to out-of-distribution (OOD) shifts. Because the quality of action directly depends on the accuracy of the predicted future visual information, OOD conditions affect both stages at once. To address this vulnerability, we propose Test-Time Training Visual Foresight VLA ($T^3$VF), a test-time training approach motivated by the observation that the predicted future image and its subsequent observation form a natural supervision pair. To further address the practical challenges that arise from indiscriminate test-time updates, we introduce an adaptive update filtering mechanism. Empirically, $T^3$VF mitigates the OOD vulnerability of VF-VLA at a modest additional inference cost, without requiring any architectural modification or auxiliary modules.
title Test-Time Training for Visual Foresight Vision-Language-Action Models
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2605.08215