VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ziakas, Christos, Russo, Alessandra |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025)
by: Romero, Angel, et al.
Published: (2025)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026)
by: Babu, Abhijith, et al.
Published: (2026)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026)
by: Liu, Zhi
Published: (2026)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
by: Römer, Ralf, et al.
Published: (2026)
by: Römer, Ralf, et al.
Published: (2026)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
by: Gasparino, Mateus Valverde, et al.
Published: (2024)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
by: Syed, Shahram Najam, et al.
Published: (2025)
by: Syed, Shahram Najam, et al.
Published: (2025)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
by: Mansour, Jad, et al.
Published: (2025)
by: Mansour, Jad, et al.
Published: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
by: Mansour, Jad, et al.
Published: (2024)
by: Mansour, Jad, et al.
Published: (2024)
HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos
by: Wang, Zhi, et al.
Published: (2026)
by: Wang, Zhi, et al.
Published: (2026)
Pointing-Guided Target Estimation via Transformer-Based Attention
by: Müller, Luca, et al.
Published: (2025)
by: Müller, Luca, et al.
Published: (2025)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
by: Aida, Adriana, et al.
Published: (2026)
by: Aida, Adriana, et al.
Published: (2026)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
by: Ahmadi, Ehsan, et al.
Published: (2024)
by: Ahmadi, Ehsan, et al.
Published: (2024)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
by: Hu, X., et al.
Published: (2025)
by: Hu, X., et al.
Published: (2025)
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
by: Li, Xinting, et al.
Published: (2023)
by: Li, Xinting, et al.
Published: (2023)
Convolutional Model Trees
by: Armstrong, William Ward, et al.
Published: (2025)
by: Armstrong, William Ward, et al.
Published: (2025)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
by: Ai, Bo, et al.
Published: (2024)
by: Ai, Bo, et al.
Published: (2024)
VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
by: Mercat, Jean, et al.
Published: (2026)
by: Mercat, Jean, et al.
Published: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
by: Seo, Huichan, et al.
Published: (2025)
by: Seo, Huichan, et al.
Published: (2025)
Contextual Graph Representations for Task-Driven 3D Perception and Planning
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D
by: Arnaud, Sergio, et al.
Published: (2025)
by: Arnaud, Sergio, et al.
Published: (2025)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
by: Agia, Christopher, et al.
Published: (2024)
by: Agia, Christopher, et al.
Published: (2024)
Failure Prediction at Runtime for Generative Robot Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
The Impact of 2D Segmentation Backbones on Point Cloud Predictions Using 4D Radar
by: Muckelroy III, William, et al.
Published: (2025)
by: Muckelroy III, William, et al.
Published: (2025)
Goal-Based Vision-Language Driving
by: Patapati, Santosh, et al.
Published: (2025)
by: Patapati, Santosh, et al.
Published: (2025)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
by: Endo, Masafumi, et al.
Published: (2024)
by: Endo, Masafumi, et al.
Published: (2024)
Single-Shot Metric Depth from Focused Plenoptic Cameras
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
by: Nguyen, Viet Dung, et al.
Published: (2026)
by: Nguyen, Viet Dung, et al.
Published: (2026)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
by: Mahdian, Navid, et al.
Published: (2024)
by: Mahdian, Navid, et al.
Published: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
by: Ferenczi, Bryce, et al.
Published: (2023)
by: Ferenczi, Bryce, et al.
Published: (2023)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
by: Tourani, Ali, et al.
Published: (2023)
by: Tourani, Ali, et al.
Published: (2023)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
by: Wu, Jason, et al.
Published: (2026)
by: Wu, Jason, et al.
Published: (2026)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
by: Mehta, Vinit, et al.
Published: (2025)
by: Mehta, Vinit, et al.
Published: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
by: Kang, Li, et al.
Published: (2026)
by: Kang, Li, et al.
Published: (2026)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
by: Bartkowiak, Patryk, et al.
Published: (2026)
by: Bartkowiak, Patryk, et al.
Published: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Similar Items
-
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
by: Romero, Angel, et al.
Published: (2025) -
Closed-Loop Neural Activation Control in Vision-Language-Action Models
by: Babu, Abhijith, et al.
Published: (2026) -
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
by: Liu, Zhi
Published: (2026) -
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
by: Römer, Ralf, et al.
Published: (2026) -
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
by: Gasparino, Mateus Valverde, et al.
Published: (2024)