SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Haowen, Yao, Shaoxiong, Chen, Haonan, Gao, Jiawei, Mao, Jiayuan, Huang, Jia-Bin, Du, Yilun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
von: Lee, Seungjae, et al.
Veröffentlicht: (2025)
von: Lee, Seungjae, et al.
Veröffentlicht: (2025)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
von: Dai, Tingjun, et al.
Veröffentlicht: (2026)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
von: Gao, Chongkai, et al.
Veröffentlicht: (2025)
von: Gao, Chongkai, et al.
Veröffentlicht: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
von: Ma, Guoqing, et al.
Veröffentlicht: (2026)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
von: Hui, Chenyu, et al.
Veröffentlicht: (2026)
von: Hui, Chenyu, et al.
Veröffentlicht: (2026)
A Survey on Efficient Vision-Language-Action Models
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
von: Lv, Qi, et al.
Veröffentlicht: (2025)
von: Lv, Qi, et al.
Veröffentlicht: (2025)
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
von: Lin, Tao, et al.
Veröffentlicht: (2026)
von: Lin, Tao, et al.
Veröffentlicht: (2026)
Unified Vision-Language-Action Model
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
von: Wang, Yuqi, et al.
Veröffentlicht: (2025)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
von: Li, Chenyang, et al.
Veröffentlicht: (2026)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
3D-VLA: A 3D Vision-Language-Action Generative World Model
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2024)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
von: Huang, Wenhui, et al.
Veröffentlicht: (2026)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
von: Ding, Pengxiang, et al.
Veröffentlicht: (2023)
AdaWorld: Learning Adaptable World Models with Latent Actions
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
von: Gao, Shenyuan, et al.
Veröffentlicht: (2025)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
von: Li, Qiwei, et al.
Veröffentlicht: (2026)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyao, et al.
Veröffentlicht: (2025)
EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
von: Bai, Zechen, et al.
Veröffentlicht: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
von: Tang, Zuojin, et al.
Veröffentlicht: (2026)
StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
von: Shi, Yiran, et al.
Veröffentlicht: (2026)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
von: Huang, Huang, et al.
Veröffentlicht: (2025)
von: Huang, Huang, et al.
Veröffentlicht: (2025)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2025)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
von: Fu, Yankai, et al.
Veröffentlicht: (2025)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
von: Feng, Yixu, et al.
Veröffentlicht: (2026)
von: Feng, Yixu, et al.
Veröffentlicht: (2026)
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
von: Lin, Tao, et al.
Veröffentlicht: (2025)
von: Lin, Tao, et al.
Veröffentlicht: (2025)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models
von: Yang, Jiabing, et al.
Veröffentlicht: (2026)
von: Yang, Jiabing, et al.
Veröffentlicht: (2026)
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
von: Li, Haosheng, et al.
Veröffentlicht: (2026)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
von: Chen, Shizhe, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
von: Lee, Seungjae, et al.
Veröffentlicht: (2025) -
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025) -
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
von: Dai, Tingjun, et al.
Veröffentlicht: (2026) -
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025) -
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)