From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Yihan, Li, Haoyang, Li, Yang, Shen, Haitao, Zhao, Yihan, Shao, Chao, Zhang, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
von: Zhao, Chen, et al.
Veröffentlicht: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
von: Wang, Yating, et al.
Veröffentlicht: (2025)
von: Wang, Yating, et al.
Veröffentlicht: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
von: Liu, Chenghao, et al.
Veröffentlicht: (2025)
ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
von: Ye, Wencheng, et al.
Veröffentlicht: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
von: Li, Yajie, et al.
Veröffentlicht: (2026)
von: Li, Yajie, et al.
Veröffentlicht: (2026)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
von: Sun, Jingwen, et al.
Veröffentlicht: (2026)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
von: Nie, Dujun, et al.
Veröffentlicht: (2026)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
LLaDA-VLA: Vision Language Diffusion Action Models
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
von: Wen, Yuqing, et al.
Veröffentlicht: (2025)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2025)
von: Din, Muhayy Ud, et al.
Veröffentlicht: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
von: Zhao, Baining, et al.
Veröffentlicht: (2026)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
von: Yang, Jiabing, et al.
Veröffentlicht: (2026)
von: Yang, Jiabing, et al.
Veröffentlicht: (2026)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
von: Zhang, Borong, et al.
Veröffentlicht: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
von: Song, Wenxuan, et al.
Veröffentlicht: (2026)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
von: Shao, Rui, et al.
Veröffentlicht: (2025)
von: Shao, Rui, et al.
Veröffentlicht: (2025)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
von: Wang, Linbo, et al.
Veröffentlicht: (2026)
von: Wang, Linbo, et al.
Veröffentlicht: (2026)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
von: Yang, Yurou, et al.
Veröffentlicht: (2026)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
von: Li, Zhuohao, et al.
Veröffentlicht: (2026)
CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
von: Lin, Tao, et al.
Veröffentlicht: (2026)
von: Lin, Tao, et al.
Veröffentlicht: (2026)
GigaBrain-0: A World Model-Powered Vision-Language-Action Model
von: GigaBrain Team, et al.
Veröffentlicht: (2025)
von: GigaBrain Team, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
von: Zhao, Chen, et al.
Veröffentlicht: (2026) -
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026) -
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025) -
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025) -
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)