Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guruprasad, Pranav, Wang, Yangyue, Chowdhury, Sudipta, Sikka, Harshvardhan, Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2024)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2024)
Benchmarking the Generality of Vision-Language-Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
von: Wang, Yangyue, et al.
Veröffentlicht: (2026)
Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
von: Xu, Haochuan, et al.
Veröffentlicht: (2025)
Tactile Modality Fusion for Vision-Language-Action Models
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
von: Morissette, Charlotte, et al.
Veröffentlicht: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhixuan, et al.
Veröffentlicht: (2025)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
von: Zhang, Zezhou, et al.
Veröffentlicht: (2026)
Improving Vision-Language-Action Model with Online Reinforcement Learning
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
Hybrid Training for Vision-Language-Action Models
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
von: Mazzaglia, Pietro, et al.
Veröffentlicht: (2025)
Test-Time Training for Visual Foresight Vision-Language-Action Models
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
von: Park, Sangwu, et al.
Veröffentlicht: (2026)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengshen, et al.
Veröffentlicht: (2025)
DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
Vision-Language Models Unlock Task-Centric Latent Actions
von: Nikulin, Alexander, et al.
Veröffentlicht: (2026)
von: Nikulin, Alexander, et al.
Veröffentlicht: (2026)
PointVLA: Injecting the 3D World into Vision-Language-Action Models
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
von: Li, Chengmeng, et al.
Veröffentlicht: (2025)
A Survey on Efficient Vision-Language-Action Models
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
Interactive Post-Training for Vision-Language-Action Models
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
von: Tan, Shuhan, et al.
Veröffentlicht: (2025)
Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks
von: Shandilya, Utkarsh, et al.
Veröffentlicht: (2025)
von: Shandilya, Utkarsh, et al.
Veröffentlicht: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
Spatio-Temporal LLM: Reasoning about Environments and Actions
von: Zheng, Haozhen, et al.
Veröffentlicht: (2025)
von: Zheng, Haozhen, et al.
Veröffentlicht: (2025)
The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling
von: Shiba, Takuya
Veröffentlicht: (2026)
von: Shiba, Takuya
Veröffentlicht: (2026)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
Taylor Videos for Action Recognition
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Action-Inspired Generative Models
von: A., Eshwar R., et al.
Veröffentlicht: (2026)
von: A., Eshwar R., et al.
Veröffentlicht: (2026)
Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
von: Taherin, Amir, et al.
Veröffentlicht: (2025)
von: Taherin, Amir, et al.
Veröffentlicht: (2025)
AEGIS: Anchor-Enforced Gradient Isolation for Knowledge-Preserving Vision-Language-Action Fine-Tuning
von: Singh, Guransh
Veröffentlicht: (2026)
von: Singh, Guransh
Veröffentlicht: (2026)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
von: Li, Qixiu, et al.
Veröffentlicht: (2024)
FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
von: An, Xinyuan, et al.
Veröffentlicht: (2026)
von: An, Xinyuan, et al.
Veröffentlicht: (2026)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2024)
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2024)
Progressive Compositionality in Text-to-Image Generative Models
von: Han, Evans Xu, et al.
Veröffentlicht: (2024)
von: Han, Evans Xu, et al.
Veröffentlicht: (2024)
A Vision for Multisensory Intelligence: Sensing, Science, and Synergy
von: Liang, Paul Pu
Veröffentlicht: (2026)
von: Liang, Paul Pu
Veröffentlicht: (2026)
Benchmarking Sensitivity of Continual Graph Learning for Skeleton-Based Action Recognition
von: Wei, Wei, et al.
Veröffentlicht: (2024)
von: Wei, Wei, et al.
Veröffentlicht: (2024)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
von: Xu, Siyu, et al.
Veröffentlicht: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025) -
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2024) -
Benchmarking the Generality of Vision-Language-Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025) -
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
von: Wang, Yangyue, et al.
Veröffentlicht: (2026) -
Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
von: Xu, Yi, et al.
Veröffentlicht: (2025)