Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Xudong, Yang, Yaoxin, Ye, Peng, Zheng, Jialin, Bai, Bizhe, Wang, Xinyi, Hao, Jia, Chen, Tao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
von: Singhi, Nishad, et al.
Veröffentlicht: (2026)
von: Singhi, Nishad, et al.
Veröffentlicht: (2026)
Local Information Matters: Inference Acceleration For Grounded Conversation Generation Models Through Adaptive Local-Aware Token Pruning
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025)
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
von: Izzo, Riccardo Andrea, et al.
Veröffentlicht: (2026)
von: Izzo, Riccardo Andrea, et al.
Veröffentlicht: (2026)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
von: Tan, Xudong, et al.
Veröffentlicht: (2025)
Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making
von: Wan, Xu, et al.
Veröffentlicht: (2025)
von: Wan, Xu, et al.
Veröffentlicht: (2025)
Learning to Explore with Parameter-Space Noise: A Deep Dive into Parameter-Space Noise for Reinforcement Learning with Verifiable Rewards
von: Bai, Bizhe, et al.
Veröffentlicht: (2026)
von: Bai, Bizhe, et al.
Veröffentlicht: (2026)
M-GRPO: Stabilizing Self-Supervised Reinforcement Learning for Large Language Models with Momentum-Anchored Policy Optimization
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
von: Phan, Hoang, et al.
Veröffentlicht: (2025)
von: Phan, Hoang, et al.
Veröffentlicht: (2025)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
von: Qin, Shengling, et al.
Veröffentlicht: (2025)
Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
von: Ye, Yifan, et al.
Veröffentlicht: (2025)
von: Ye, Yifan, et al.
Veröffentlicht: (2025)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2026)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2026)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
von: Pertsch, Karl, et al.
Veröffentlicht: (2025)
von: Pertsch, Karl, et al.
Veröffentlicht: (2025)
Think Twice Before You Act: Improving Inverse Problem Solving With MCMC
von: Zhu, Yaxuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yaxuan, et al.
Veröffentlicht: (2024)
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
von: Ling, Yiran, et al.
Veröffentlicht: (2026)
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
von: Li, Ye, et al.
Veröffentlicht: (2026)
von: Li, Ye, et al.
Veröffentlicht: (2026)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
von: Tang, Fei, et al.
Veröffentlicht: (2025)
von: Tang, Fei, et al.
Veröffentlicht: (2025)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
von: Chen, Guangyan, et al.
Veröffentlicht: (2025)
von: Chen, Guangyan, et al.
Veröffentlicht: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyue, et al.
Veröffentlicht: (2025)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
von: Liu, Shuming, et al.
Veröffentlicht: (2026)
von: Liu, Shuming, et al.
Veröffentlicht: (2026)
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
von: Liang, Yuanchang, et al.
Veröffentlicht: (2026)
ActionCodec: What Makes for Good Action Tokenizers
von: Dong, Zibin, et al.
Veröffentlicht: (2026)
von: Dong, Zibin, et al.
Veröffentlicht: (2026)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
von: Shiba, Takuya
Veröffentlicht: (2026)
von: Shiba, Takuya
Veröffentlicht: (2026)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
von: Xiong, Zheng, et al.
Veröffentlicht: (2025)
von: Xiong, Zheng, et al.
Veröffentlicht: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
von: Peng, Ziqiao, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making
von: Son, Young-Chae, et al.
Veröffentlicht: (2026)
von: Son, Young-Chae, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
von: Singhi, Nishad, et al.
Veröffentlicht: (2026) -
Local Information Matters: Inference Acceleration For Grounded Conversation Generation Models Through Adaptive Local-Aware Token Pruning
von: Bai, Bizhe, et al.
Veröffentlicht: (2025) -
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
von: Yang, Yaoxin, et al.
Veröffentlicht: (2025) -
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
von: Izzo, Riccardo Andrea, et al.
Veröffentlicht: (2026) -
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
von: Tan, Xudong, et al.
Veröffentlicht: (2025)