Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guruprasad, Pranav, Sikka, Harshvardhan, Song, Jaewoo, Wang, Yangyue, Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
Improving Vision-Language-Action Model with Online Reinforcement Learning
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
Benchmarking the Generality of Vision-Language-Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
Tactile Modality Fusion for Vision-Language-Action Models
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
A Survey on Efficient Vision-Language-Action Models
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
von: Wang, Beichen, et al.
Veröffentlicht: (2024)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
von: An, Xinyuan, et al.
Veröffentlicht: (2026)
von: An, Xinyuan, et al.
Veröffentlicht: (2026)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2025)
von: Kawaharazuka, Kento, et al.
Veröffentlicht: (2025)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
Pedestrian Trajectory Prediction with Missing Data: Datasets, Imputation, and Benchmarking
von: Chib, Pranav Singh, et al.
Veröffentlicht: (2024)
von: Chib, Pranav Singh, et al.
Veröffentlicht: (2024)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
von: Shiba, Takuya
Veröffentlicht: (2026)
von: Shiba, Takuya
Veröffentlicht: (2026)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
von: Zhao, Yanpeng, et al.
Veröffentlicht: (2026)
von: Zhao, Yanpeng, et al.
Veröffentlicht: (2026)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
von: Zhou, Xunlan, et al.
Veröffentlicht: (2026)
von: Zhou, Xunlan, et al.
Veröffentlicht: (2026)
Robustness Evaluation of Machine Learning Models for Robot Arm Action Recognition in Noisy Environments
von: Motamedi, Elaheh, et al.
Veröffentlicht: (2024)
von: Motamedi, Elaheh, et al.
Veröffentlicht: (2024)
Hybrid Training for Vision-Language-Action Models
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
von: Zhang, Wenbo, et al.
Veröffentlicht: (2025)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
von: Cheang, Chi-Lam, et al.
Veröffentlicht: (2024)
von: Cheang, Chi-Lam, et al.
Veröffentlicht: (2024)
PerAct2: Benchmarking and Learning for Robotic Bimanual Manipulation Tasks
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
von: Grotz, Markus, et al.
Veröffentlicht: (2024)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
Interactive Post-Training for Vision-Language-Action Models
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
von: Luo, Hao, et al.
Veröffentlicht: (2025)
von: Luo, Hao, et al.
Veröffentlicht: (2025)
GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
von: Ze, Yanjie, et al.
Veröffentlicht: (2023)
von: Ze, Yanjie, et al.
Veröffentlicht: (2023)
Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
von: Taherin, Amir, et al.
Veröffentlicht: (2025)
von: Taherin, Amir, et al.
Veröffentlicht: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
von: Govind, Manish Kumar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025) -
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025) -
Improving Vision-Language-Action Model with Online Reinforcement Learning
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025) -
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025) -
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)