DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Siyuan, Wang, Tianshi, Li, Fengling, Zhu, Lei, Shen, Heng Tao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023)
by: Wang, Tianshi, et al.
Published: (2023)
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
by: Tan, Wentao, et al.
Published: (2025)
by: Tan, Wentao, et al.
Published: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
by: Zhai, Jiajun, et al.
Published: (2026)
by: Zhai, Jiajun, et al.
Published: (2026)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
by: Yan, Feng, et al.
Published: (2024)
by: Yan, Feng, et al.
Published: (2024)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025)
by: Sun, Qiao, et al.
Published: (2025)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
by: Pueyo, Pablo, et al.
Published: (2025)
by: Pueyo, Pablo, et al.
Published: (2025)
Bringing Robots Home: The Rise of AI Robots in Consumer Electronics
by: Dong, Haiwei, et al.
Published: (2024)
by: Dong, Haiwei, et al.
Published: (2024)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
by: Hsieh, Po-Yu, et al.
Published: (2024)
by: Hsieh, Po-Yu, et al.
Published: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Flight Patterns for Swarms of Drones
by: Zhu, Shuqin, et al.
Published: (2024)
by: Zhu, Shuqin, et al.
Published: (2024)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
by: Xu, Junhao, et al.
Published: (2025)
by: Xu, Junhao, et al.
Published: (2025)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
by: Wei, Yiping, et al.
Published: (2023)
by: Wei, Yiping, et al.
Published: (2023)
Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
by: Fan, Congyi, et al.
Published: (2026)
by: Fan, Congyi, et al.
Published: (2026)
MotiBo: The Impact of Interactive Digital Storytelling Robots on Student Motivation through Self-Determination Theory
by: Fung, Ka Yan, et al.
Published: (2026)
by: Fung, Ka Yan, et al.
Published: (2026)
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer
by: Zhi, Heng, et al.
Published: (2026)
by: Zhi, Heng, et al.
Published: (2026)
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
by: Li, Shenshen, et al.
Published: (2025)
by: Li, Shenshen, et al.
Published: (2025)
U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon Tasks
by: Feng, Tongtong, et al.
Published: (2025)
by: Feng, Tongtong, et al.
Published: (2025)
Teaching Physical Awareness to LLMs through Sounds
by: Wang, Weiguo, et al.
Published: (2025)
by: Wang, Weiguo, et al.
Published: (2025)
SOP: A Scalable Online Post-Training System for Vision-Language-Action Models
by: Pan, Mingjie, et al.
Published: (2026)
by: Pan, Mingjie, et al.
Published: (2026)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
by: Liu, Yanbaihui, et al.
Published: (2024)
by: Liu, Yanbaihui, et al.
Published: (2024)
TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings
by: Wasi, Azmine Toushik, et al.
Published: (2026)
by: Wasi, Azmine Toushik, et al.
Published: (2026)
WoW: Towards a World omniscient World model Through Embodied Interaction
by: Chi, Xiaowei, et al.
Published: (2025)
by: Chi, Xiaowei, et al.
Published: (2025)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
HR-INR: Continuous Space-Time Video Super-Resolution via Event Camera
by: Lu, Yunfan, et al.
Published: (2024)
by: Lu, Yunfan, et al.
Published: (2024)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
by: Qian, Kangan, et al.
Published: (2026)
by: Qian, Kangan, et al.
Published: (2026)
EQ-TAA: Equivariant Traffic Accident Anticipation via Diffusion-Based Accident Video Synthesis
by: Fang, Jianwu, et al.
Published: (2025)
by: Fang, Jianwu, et al.
Published: (2025)
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
EVA: An Embodied World Model for Future Video Anticipation
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Voxel-GS: Quantized Scaffold Gaussian Splatting Compression with Run-Length Coding
by: Fu, Chunyang, et al.
Published: (2025)
by: Fu, Chunyang, et al.
Published: (2025)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios
by: Peng, Kunyu, et al.
Published: (2025)
by: Peng, Kunyu, et al.
Published: (2025)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
by: Guo, Ziang, et al.
Published: (2026)
by: Guo, Ziang, et al.
Published: (2026)
Similar Items
-
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026) -
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
by: Wang, Tianshi, et al.
Published: (2023) -
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
by: Tan, Wentao, et al.
Published: (2025) -
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
by: Ye, Wencheng, et al.
Published: (2025) -
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
by: Zhai, Jiajun, et al.
Published: (2026)