Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Yifan, Ma, Jiaqi, Cen, Jun, Lu, Zhihe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
by: Li, Yuquan, et al.
Published: (2026)
by: Li, Yuquan, et al.
Published: (2026)
Self-evolved Imitation Learning in Simulated World
by: Ye, Yifan, et al.
Published: (2025)
by: Ye, Yifan, et al.
Published: (2025)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
by: Fang, Zhou, et al.
Published: (2026)
by: Fang, Zhou, et al.
Published: (2026)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
by: Shiba, Takuya
Published: (2026)
by: Shiba, Takuya
Published: (2026)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
by: Liu, Yicheng, et al.
Published: (2025)
by: Liu, Yicheng, et al.
Published: (2025)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
by: Xu, Charles, et al.
Published: (2026)
by: Xu, Charles, et al.
Published: (2026)
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
by: Koo, Jiyeon, et al.
Published: (2025)
by: Koo, Jiyeon, et al.
Published: (2025)
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance
by: Park, Jongwoo, et al.
Published: (2026)
by: Park, Jongwoo, et al.
Published: (2026)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent
by: Fu, Yuxia, et al.
Published: (2025)
by: Fu, Yuxia, et al.
Published: (2025)
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026)
by: Dong, Zibin, et al.
Published: (2026)
VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer
by: Hu, Songqiao, et al.
Published: (2025)
by: Hu, Songqiao, et al.
Published: (2025)
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
by: Hu, Yanbin, et al.
Published: (2026)
by: Hu, Yanbin, et al.
Published: (2026)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
by: Li, Chenyang, et al.
Published: (2026)
by: Li, Chenyang, et al.
Published: (2026)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
by: Li, Haosheng, et al.
Published: (2026)
by: Li, Haosheng, et al.
Published: (2026)
ETA-VLA: Efficient Token Adaptation via Temporal Fusion and Intra-LLM Sparsification for Vision-Language-Action Models
by: Wang, Yiru, et al.
Published: (2026)
by: Wang, Yiru, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
Action Tokenizer Matters in In-Context Imitation Learning
by: Vuong, An Dinh, et al.
Published: (2025)
by: Vuong, An Dinh, et al.
Published: (2025)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
OAT: Ordered Action Tokenization
by: Liu, Chaoqi, et al.
Published: (2026)
by: Liu, Chaoqi, et al.
Published: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
RynnVLA-002: A Unified Vision-Language-Action and World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
by: Huang, Zhiyu, et al.
Published: (2026)
by: Huang, Zhiyu, et al.
Published: (2026)
BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
by: Ma, Xiaoyu, et al.
Published: (2025)
by: Ma, Xiaoyu, et al.
Published: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
by: Sarowar, Md Selim, et al.
Published: (2026)
by: Sarowar, Md Selim, et al.
Published: (2026)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
Grid-based Fast and Structural Visual Odometry
by: Zhihe, Zhang
Published: (2024)
by: Zhihe, Zhang
Published: (2024)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
by: Yadav, Yajat, et al.
Published: (2025)
by: Yadav, Yajat, et al.
Published: (2025)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation
by: Wang, Qitong, et al.
Published: (2026)
by: Wang, Qitong, et al.
Published: (2026)
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
by: Chen, Yuxuan, et al.
Published: (2025)
by: Chen, Yuxuan, et al.
Published: (2025)
Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers
by: Chen, Yutian, et al.
Published: (2025)
by: Chen, Yutian, et al.
Published: (2025)
RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models
by: Luo, Jingzhou, et al.
Published: (2026)
by: Luo, Jingzhou, et al.
Published: (2026)
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
by: Fateh, Fawad Javed, et al.
Published: (2026)
by: Fateh, Fawad Javed, et al.
Published: (2026)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
by: Huang, Helong, et al.
Published: (2025)
by: Huang, Helong, et al.
Published: (2025)
Similar Items
-
DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference
by: Li, Yuquan, et al.
Published: (2026) -
Self-evolved Imitation Learning in Simulated World
by: Ye, Yifan, et al.
Published: (2025) -
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025) -
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025) -
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
by: Fang, Zhou, et al.
Published: (2026)