Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Juntao, Ye, Feiyang, Zhang, Jing, Qian, Wenjing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
by: Xiao, Lei, et al.
Published: (2025)
by: Xiao, Lei, et al.
Published: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
by: Almuzairee, Abdulaziz, et al.
Published: (2025)
by: Almuzairee, Abdulaziz, et al.
Published: (2025)
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024)
by: Zheng, Jinliang, et al.
Published: (2024)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
by: Zhou, Zhongyi, et al.
Published: (2025)
by: Zhou, Zhongyi, et al.
Published: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
by: Bi, Hongzhe, et al.
Published: (2025)
by: Bi, Hongzhe, et al.
Published: (2025)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation
by: Ni, Zehao, et al.
Published: (2025)
by: Ni, Zehao, et al.
Published: (2025)
Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing
by: Yuan, Ying, et al.
Published: (2023)
by: Yuan, Ying, et al.
Published: (2023)
Leveraging Locality to Boost Sample Efficiency in Robotic Manipulation
by: Zhang, Tong, et al.
Published: (2024)
by: Zhang, Tong, et al.
Published: (2024)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
by: Han, ByungOk, et al.
Published: (2024)
by: Han, ByungOk, et al.
Published: (2024)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
by: Cheang, Chi-Lam, et al.
Published: (2024)
by: Cheang, Chi-Lam, et al.
Published: (2024)
AdaptiGraph: Material-Adaptive Graph-Based Neural Dynamics for Robotic Manipulation
by: Zhang, Kaifeng, et al.
Published: (2024)
by: Zhang, Kaifeng, et al.
Published: (2024)
Evaluating Real-World Robot Manipulation Policies in Simulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
M2CURL: Sample-Efficient Multimodal Reinforcement Learning via Self-Supervised Representation Learning for Robotic Manipulation
by: Lygerakis, Fotios, et al.
Published: (2024)
by: Lygerakis, Fotios, et al.
Published: (2024)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
by: Shiba, Takuya
Published: (2026)
by: Shiba, Takuya
Published: (2026)
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
by: Zhou, Xunlan, et al.
Published: (2026)
by: Zhou, Xunlan, et al.
Published: (2026)
Visual Whole-Body Control for Legged Loco-Manipulation
by: Liu, Minghuan, et al.
Published: (2024)
by: Liu, Minghuan, et al.
Published: (2024)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
by: Niu, Dantong, et al.
Published: (2024)
by: Niu, Dantong, et al.
Published: (2024)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
ZeroMimic: Distilling Robotic Manipulation Skills from Web Videos
by: Shi, Junyao, et al.
Published: (2025)
by: Shi, Junyao, et al.
Published: (2025)
SCENEREPLICA: Benchmarking Real-World Robot Manipulation by Creating Replicable Scenes
by: Khargonkar, Ninad, et al.
Published: (2023)
by: Khargonkar, Ninad, et al.
Published: (2023)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
by: Ivanovic, Boris, et al.
Published: (2025)
by: Ivanovic, Boris, et al.
Published: (2025)
LOTUS: Continual Imitation Learning for Robot Manipulation Through Unsupervised Skill Discovery
by: Wan, Weikang, et al.
Published: (2023)
by: Wan, Weikang, et al.
Published: (2023)
Planning-Guided Diffusion Policy Learning for Generalizable Contact-Rich Bimanual Manipulation
by: Li, Xuanlin, et al.
Published: (2024)
by: Li, Xuanlin, et al.
Published: (2024)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
by: Bauer, Dominik, et al.
Published: (2024)
by: Bauer, Dominik, et al.
Published: (2024)
GraspSplats: Efficient Manipulation with 3D Feature Splatting
by: Ji, Mazeyu, et al.
Published: (2024)
by: Ji, Mazeyu, et al.
Published: (2024)
Industrial Application of 6D Pose Estimation for Robotic Manipulation in Automotive Internal Logistics
by: Quentin, Philipp, et al.
Published: (2023)
by: Quentin, Philipp, et al.
Published: (2023)
DexSkin: High-Coverage Conformable Robotic Skin for Learning Contact-Rich Manipulation
by: Wistreich, Suzannah, et al.
Published: (2025)
by: Wistreich, Suzannah, et al.
Published: (2025)
Think Proprioceptively: Embodied Visual Reasoning for VLA Manipulation
by: Wang, Fangyuan, et al.
Published: (2026)
by: Wang, Fangyuan, et al.
Published: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
by: Guo, Heyu, et al.
Published: (2025)
by: Guo, Heyu, et al.
Published: (2025)
Similar Items
-
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
by: Xiao, Lei, et al.
Published: (2025) -
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025) -
Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
by: Almuzairee, Abdulaziz, et al.
Published: (2025) -
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024) -
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
by: Zhou, Zhongyi, et al.
Published: (2025)