RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Koo, Jiyeon, Cho, Taewan, Kang, Hyunjoon, Pyo, Eunseom, Oh, Tae Gyun, Kim, Taeryang, Choi, Andrew Jaeyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
von: Cho, Taewan, et al.
Veröffentlicht: (2026)
von: Cho, Taewan, et al.
Veröffentlicht: (2026)
Transferability of Token Usage Rights: A Design Space Analysis of Generative AI Services
von: Lee, Jaeyong, et al.
Veröffentlicht: (2026)
von: Lee, Jaeyong, et al.
Veröffentlicht: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
Mini bat organs reveal hidden viral threats
von: Hyunjoon Kim, et al.
Veröffentlicht: (2025)
von: Hyunjoon Kim, et al.
Veröffentlicht: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
IdeaBlocks: Expressing and Reusing Divergent Intents for Graphic Design Exploration using Generative AI
von: Choi, DaEun, et al.
Veröffentlicht: (2025)
von: Choi, DaEun, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
von: Ye, Angen, et al.
Veröffentlicht: (2025)
von: Ye, Angen, et al.
Veröffentlicht: (2025)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
COSMosFL: Ensemble of Small Language Models for Fault Localisation
von: Cho, Hyunjoon, et al.
Veröffentlicht: (2025)
von: Cho, Hyunjoon, et al.
Veröffentlicht: (2025)
Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection
von: Lee, Tae-Gyun, et al.
Veröffentlicht: (2026)
von: Lee, Tae-Gyun, et al.
Veröffentlicht: (2026)
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
von: Yin, Cheng, et al.
Veröffentlicht: (2025)
von: Yin, Cheng, et al.
Veröffentlicht: (2025)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
von: Tu, Ruisen, et al.
Veröffentlicht: (2026)
von: Tu, Ruisen, et al.
Veröffentlicht: (2026)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
von: Pan, Xu, et al.
Veröffentlicht: (2026)
von: Pan, Xu, et al.
Veröffentlicht: (2026)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
von: Lin, Fanqi, et al.
Veröffentlicht: (2025)
von: Lin, Fanqi, et al.
Veröffentlicht: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
von: Huang, Helong, et al.
Veröffentlicht: (2025)
von: Huang, Helong, et al.
Veröffentlicht: (2025)
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
von: Jie, Haoxiang, et al.
Veröffentlicht: (2026)
von: Jie, Haoxiang, et al.
Veröffentlicht: (2026)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
von: Xiong, Zheng, et al.
Veröffentlicht: (2025)
von: Xiong, Zheng, et al.
Veröffentlicht: (2025)
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
von: Qu, Delin, et al.
Veröffentlicht: (2025)
von: Qu, Delin, et al.
Veröffentlicht: (2025)
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
von: Choi, Hyeonbeom, et al.
Veröffentlicht: (2026)
von: Choi, Hyeonbeom, et al.
Veröffentlicht: (2026)
VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
EdgeVLA: Efficient Vision-Language-Action Models
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
CRL-VLA: Continual Vision-Language-Action Learning
von: Zeng, Qixin, et al.
Veröffentlicht: (2026)
von: Zeng, Qixin, et al.
Veröffentlicht: (2026)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action Models
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
von: Wang, Yiru, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
von: Cho, Taewan, et al.
Veröffentlicht: (2026) -
Transferability of Token Usage Rights: A Design Space Analysis of Generative AI Services
von: Lee, Jaeyong, et al.
Veröffentlicht: (2026) -
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025) -
Mini bat organs reveal hidden viral threats
von: Hyunjoon Kim, et al.
Veröffentlicht: (2025) -
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)