Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jiaming, Ye, Ke, Liu, Jiayi, Ma, Teli, Wang, Zifan, Qiu, Ronghe, Lin, Kun-Yu, Zhao, Zhilin, Liang, Junwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
by: Ye, Ke, et al.
Published: (2025)
by: Ye, Ke, et al.
Published: (2025)
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
by: Ma, Teli, et al.
Published: (2026)
by: Ma, Teli, et al.
Published: (2026)
End-to-End Humanoid Robot Safe and Comfortable Locomotion Policy
by: Wang, Zifan, et al.
Published: (2025)
by: Wang, Zifan, et al.
Published: (2025)
Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments
by: Wang, Zifan, et al.
Published: (2025)
by: Wang, Zifan, et al.
Published: (2025)
From Cognition to Precognition: A Future-Aware Framework for Social Navigation
by: Gong, Zeying, et al.
Published: (2024)
by: Gong, Zeying, et al.
Published: (2024)
EgoTraj-Bench: Towards Robust Trajectory Prediction Under Ego-view Noisy Observations
by: Liu, Jiayi, et al.
Published: (2025)
by: Liu, Jiayi, et al.
Published: (2025)
SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
by: Qiu, Dicong, et al.
Published: (2025)
by: Qiu, Dicong, et al.
Published: (2025)
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
by: Huang, Yuzhe, et al.
Published: (2026)
by: Huang, Yuzhe, et al.
Published: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
Cross-Hand Latent Representation for Vision-Language-Action Models
by: Jiang, Guangqi, et al.
Published: (2026)
by: Jiang, Guangqi, et al.
Published: (2026)
An Examination of the Compositionality of Large Generative Vision-Language Models
by: Ma, Teli, et al.
Published: (2023)
by: Ma, Teli, et al.
Published: (2023)
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
by: Hu, Yucheng, et al.
Published: (2026)
by: Hu, Yucheng, et al.
Published: (2026)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
Vision-Language Model Predictive Control for Manipulation Planning and Trajectory Generation
by: Chen, Jiaming, et al.
Published: (2025)
by: Chen, Jiaming, et al.
Published: (2025)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
by: Li, Pengteng, et al.
Published: (2026)
by: Li, Pengteng, et al.
Published: (2026)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
by: Zhang, Yuhao, et al.
Published: (2026)
by: Zhang, Yuhao, et al.
Published: (2026)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
by: Liu, Zhuoyang, et al.
Published: (2025)
by: Liu, Zhuoyang, et al.
Published: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation
by: Wang, Zhijie, et al.
Published: (2024)
by: Wang, Zhijie, et al.
Published: (2024)
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration
by: Gong, Zeying, et al.
Published: (2025)
by: Gong, Zeying, et al.
Published: (2025)
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
by: Hu, Yanbin, et al.
Published: (2026)
by: Hu, Yanbin, et al.
Published: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
by: Pei, Xiaohuan, et al.
Published: (2025)
by: Pei, Xiaohuan, et al.
Published: (2025)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
by: Song, Kun, et al.
Published: (2025)
by: Song, Kun, et al.
Published: (2025)
Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation
by: Truongcao, Keith, et al.
Published: (2026)
by: Truongcao, Keith, et al.
Published: (2026)
TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation
by: Zhang, Kaidi, et al.
Published: (2026)
by: Zhang, Kaidi, et al.
Published: (2026)
TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation
by: Sun, Jun, et al.
Published: (2026)
by: Sun, Jun, et al.
Published: (2026)
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
by: Hu, Tianshuai, et al.
Published: (2025)
by: Hu, Tianshuai, et al.
Published: (2025)
Exploring the Adversarial Vulnerabilities of Vision-Language-Action Models in Robotics
by: Wang, Taowen, et al.
Published: (2024)
by: Wang, Taowen, et al.
Published: (2024)
IA-VLA: Input Augmentation for Vision-Language-Action models in settings with semantically complex tasks
by: Hannus, Eric, et al.
Published: (2025)
by: Hannus, Eric, et al.
Published: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
by: Deng, Shengliang, et al.
Published: (2025)
by: Deng, Shengliang, et al.
Published: (2025)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
by: Peng, Xiongfeng, et al.
Published: (2026)
by: Peng, Xiongfeng, et al.
Published: (2026)
Similar Items
-
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
by: Ma, Teli, et al.
Published: (2024) -
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
by: Zhou, Jiaming, et al.
Published: (2024) -
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025) -
From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
by: Ye, Ke, et al.
Published: (2025) -
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
by: Ma, Teli, et al.
Published: (2024)