SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Hengyu, Liu, Yijiang, Du, Yuan, Du, Li, Yang, Huanrui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
von: Liu, Ziyan, et al.
Veröffentlicht: (2025)
3D-VLA: A 3D Vision-Language-Action Generative World Model
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
von: Singh, Ishika, et al.
Veröffentlicht: (2025)
VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models
von: Ranjan, Ravi, et al.
Veröffentlicht: (2026)
von: Ranjan, Ravi, et al.
Veröffentlicht: (2026)
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
von: Gao, Chongkai, et al.
Veröffentlicht: (2025)
von: Gao, Chongkai, et al.
Veröffentlicht: (2025)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
von: Shen, Boyang, et al.
Veröffentlicht: (2026)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
von: Li, Anqi, et al.
Veröffentlicht: (2025)
von: Li, Anqi, et al.
Veröffentlicht: (2025)
SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration
von: Li, Ye, et al.
Veröffentlicht: (2025)
von: Li, Ye, et al.
Veröffentlicht: (2025)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
von: Si, Shengyu, et al.
Veröffentlicht: (2026)
von: Si, Shengyu, et al.
Veröffentlicht: (2026)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
GST-VLA: Structured Gaussian Spatial Tokens for 3D Depth-Aware Vision-Language-Action Models
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2026)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games
von: Chen, Peng, et al.
Veröffentlicht: (2025)
von: Chen, Peng, et al.
Veröffentlicht: (2025)
PAT: Pruning-Aware Tuning for Large Language Models
von: Liu, Yijiang, et al.
Veröffentlicht: (2024)
von: Liu, Yijiang, et al.
Veröffentlicht: (2024)
EvoVLA: Self-Evolving Vision-Language-Action Model
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
von: Liu, Zeting, et al.
Veröffentlicht: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zheyuan, et al.
Veröffentlicht: (2026)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
von: Sun, Guoheng, et al.
Veröffentlicht: (2026)
von: Sun, Guoheng, et al.
Veröffentlicht: (2026)
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
von: Gao, Mingjian, et al.
Veröffentlicht: (2026)
von: Gao, Mingjian, et al.
Veröffentlicht: (2026)
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
VLA-IAP: Training-Free Visual Token Pruning via Interaction Alignment for Vision-Language-Action Models
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
von: Cheng, Jintao, et al.
Veröffentlicht: (2026)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
von: Jiang, Anqing, et al.
Veröffentlicht: (2025)
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
von: Community, StarVLA
Veröffentlicht: (2026)
von: Community, StarVLA
Veröffentlicht: (2026)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
von: Chen, Xinyi, et al.
Veröffentlicht: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
von: Guo, Wenxuan, et al.
Veröffentlicht: (2026)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
von: Du, Yiyang, et al.
Veröffentlicht: (2026)
von: Du, Yiyang, et al.
Veröffentlicht: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
von: Wang, Hanzhen, et al.
Veröffentlicht: (2025) -
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026) -
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026) -
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025) -
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)