AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Wenda, Wang, Tianshi, Li, Fengling, Li, Jingjing, Zhu, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
von: Wang, Tianshi, et al.
Veröffentlicht: (2023)
von: Wang, Tianshi, et al.
Veröffentlicht: (2023)
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
von: Yan, Feng, et al.
Veröffentlicht: (2024)
von: Yan, Feng, et al.
Veröffentlicht: (2024)
A Multimedia Framework for Continuum Robots: Systematic, Computational, and Control Perspectives
von: Hsieh, Po-Yu, et al.
Veröffentlicht: (2024)
von: Hsieh, Po-Yu, et al.
Veröffentlicht: (2024)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
von: Wen, Junjie, et al.
Veröffentlicht: (2024)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
von: Wei, Xiangyi, et al.
Veröffentlicht: (2025)
von: Wei, Xiangyi, et al.
Veröffentlicht: (2025)
BLM$_1$: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
von: Tan, Wentao, et al.
Veröffentlicht: (2025)
von: Tan, Wentao, et al.
Veröffentlicht: (2025)
DySL-VLA: Efficient Vision-Language-Action Model Inference via Dynamic-Static Layer-Skipping for Robot Manipulation
von: Yang, Zebin, et al.
Veröffentlicht: (2026)
von: Yang, Zebin, et al.
Veröffentlicht: (2026)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
von: Zhang, Kaidi, et al.
Veröffentlicht: (2026)
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
von: Fan, Yiguo, et al.
Veröffentlicht: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
von: Li, Runhao, et al.
Veröffentlicht: (2025)
von: Li, Runhao, et al.
Veröffentlicht: (2025)
Bringing Robots Home: The Rise of AI Robots in Consumer Electronics
von: Dong, Haiwei, et al.
Veröffentlicht: (2024)
von: Dong, Haiwei, et al.
Veröffentlicht: (2024)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
von: Sun, Jianli, et al.
Veröffentlicht: (2026)
von: Sun, Jianli, et al.
Veröffentlicht: (2026)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
Identity-Aware Vision-Language Model for Explainable Face Forgery Detection
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
von: Sun, Qiao, et al.
Veröffentlicht: (2025)
von: Sun, Qiao, et al.
Veröffentlicht: (2025)
CineWild: Balancing Art and Robotics for Ethical Wildlife Documentary Filmmaking
von: Pueyo, Pablo, et al.
Veröffentlicht: (2025)
von: Pueyo, Pablo, et al.
Veröffentlicht: (2025)
Bi-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Dexterous Manipulations
von: Gbagbe, Koffivi Fidèle, et al.
Veröffentlicht: (2024)
von: Gbagbe, Koffivi Fidèle, et al.
Veröffentlicht: (2024)
SG-VLA: Learning Spatially-Grounded Vision-Language-Action Models for Mobile Manipulation
von: Tu, Ruisen, et al.
Veröffentlicht: (2026)
von: Tu, Ruisen, et al.
Veröffentlicht: (2026)
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
von: Cai, Junhao, et al.
Veröffentlicht: (2026)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
von: Miao, Cui, et al.
Veröffentlicht: (2025)
von: Miao, Cui, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
von: Jang, Huiwon, et al.
Veröffentlicht: (2025)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2026)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2026)
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
Shake-VLA: Vision-Language-Action Model-Based System for Bimanual Robotic Manipulations and Liquid Mixing
von: Khan, Muhamamd Haris, et al.
Veröffentlicht: (2025)
von: Khan, Muhamamd Haris, et al.
Veröffentlicht: (2025)
Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
von: Guo, Ziang, et al.
Veröffentlicht: (2026)
von: Guo, Ziang, et al.
Veröffentlicht: (2026)
FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
von: Zhao, Ruiteng, et al.
Veröffentlicht: (2026)
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
von: Li, Ye, et al.
Veröffentlicht: (2026)
von: Li, Ye, et al.
Veröffentlicht: (2026)
Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision
von: Wei, Yiping, et al.
Veröffentlicht: (2023)
von: Wei, Yiping, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
von: Xu, Siyuan, et al.
Veröffentlicht: (2026) -
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
von: Zhai, Jiajun, et al.
Veröffentlicht: (2026) -
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025) -
Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
von: Wang, Tianshi, et al.
Veröffentlicht: (2023) -
RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation
von: Yan, Feng, et al.
Veröffentlicht: (2024)