Hybrid Training for Vision-Language-Action Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mazzaglia, Pietro, Sancaktar, Cansu, Peschl, Markus, Dijkman, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
Information-driven Affordance Discovery for Efficient Robotic Manipulation
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
by: Peschl, Markus, et al.
Published: (2025)
by: Peschl, Markus, et al.
Published: (2025)
Redundancy-aware Action Spaces for Robot Learning
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
Interactive Post-Training for Vision-Language-Action Models
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
GenRL: Multimodal-foundation world models for generalization in embodied agents
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
by: Zhang, Zhengshen, et al.
Published: (2025)
by: Zhang, Zhengshen, et al.
Published: (2025)
Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
by: Kim, Moo Jin, et al.
Published: (2025)
by: Kim, Moo Jin, et al.
Published: (2025)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
by: Zhao, Qingqing, et al.
Published: (2025)
by: Zhao, Qingqing, et al.
Published: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025)
by: Kawaharazuka, Kento, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models
by: Chen, Yiye, et al.
Published: (2026)
by: Chen, Yiye, et al.
Published: (2026)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
by: Hou, Yuchen, et al.
Published: (2026)
by: Hou, Yuchen, et al.
Published: (2026)
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
by: Wang, Beichen, et al.
Published: (2024)
by: Wang, Beichen, et al.
Published: (2024)
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation
by: Sautenkov, Oleg, et al.
Published: (2025)
by: Sautenkov, Oleg, et al.
Published: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
by: HU, Haibo, et al.
Published: (2025)
by: HU, Haibo, et al.
Published: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
by: Huang, Chi-Pin, et al.
Published: (2026)
by: Huang, Chi-Pin, et al.
Published: (2026)
Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
by: Larchenko, Ilia, et al.
Published: (2025)
by: Larchenko, Ilia, et al.
Published: (2025)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
by: Babey, Nicholas, et al.
Published: (2025)
by: Babey, Nicholas, et al.
Published: (2025)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
by: Vasa, Santosh, et al.
Published: (2025)
by: Vasa, Santosh, et al.
Published: (2025)
MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
Cross-Platform Scaling of Vision-Language-Action Models from Edge to Cloud GPUs
by: Taherin, Amir, et al.
Published: (2025)
by: Taherin, Amir, et al.
Published: (2025)
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation
by: Rajabi, Navid, et al.
Published: (2025)
by: Rajabi, Navid, et al.
Published: (2025)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
by: Goyal, Divyanshu, et al.
Published: (2026)
by: Goyal, Divyanshu, et al.
Published: (2026)
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System
by: Karadag, Ozge Oztimur
Published: (2023)
by: Karadag, Ozge Oztimur
Published: (2023)
From Pixels to Predicates: Learning Symbolic World Models via Pretrained Vision-Language Models
by: Athalye, Ashay, et al.
Published: (2024)
by: Athalye, Ashay, et al.
Published: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
by: Gao, Haoxiang, et al.
Published: (2025)
by: Gao, Haoxiang, et al.
Published: (2025)
Addressing the Waypoint-Action Gap in End-to-End Autonomous Driving via Vehicle Motion Models
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Similar Items
-
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025) -
Information-driven Affordance Discovery for Efficient Robotic Manipulation
by: Mazzaglia, Pietro, et al.
Published: (2024) -
From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
by: Peschl, Markus, et al.
Published: (2025) -
Redundancy-aware Action Spaces for Robot Learning
by: Mazzaglia, Pietro, et al.
Published: (2024) -
Interactive Post-Training for Vision-Language-Action Models
by: Tan, Shuhan, et al.
Published: (2025)