Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mingyu, Huang, Zheng, Lin, Xiaoyi, Zhu, Muzhi, Zhao, Canyu, Du, Zongze, Wang, Yating, Zhu, Haoyi, Chen, Hao, Shen, Chunhua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
by: Huang, Zheng, et al.
Published: (2025)
by: Huang, Zheng, et al.
Published: (2025)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
InterPreT: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning
by: Han, Muzhi, et al.
Published: (2024)
by: Han, Muzhi, et al.
Published: (2024)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning
by: Yang, Jiange, et al.
Published: (2024)
by: Yang, Jiange, et al.
Published: (2024)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
by: Zhao, Canyu, et al.
Published: (2025)
by: Zhao, Canyu, et al.
Published: (2025)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
by: Yang, Jiange, et al.
Published: (2025)
by: Yang, Jiange, et al.
Published: (2025)
Bridging VLM and KMP: Enabling Fine-grained robotic manipulation via Semantic Keypoints Representation
by: Zhu, Junjie, et al.
Published: (2025)
by: Zhu, Junjie, et al.
Published: (2025)
FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning
by: Yokoyama, Naoki, et al.
Published: (2025)
by: Yokoyama, Naoki, et al.
Published: (2025)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
by: Fan, Yiguo, et al.
Published: (2025)
by: Fan, Yiguo, et al.
Published: (2025)
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
by: Li, Puhao, et al.
Published: (2024)
by: Li, Puhao, et al.
Published: (2024)
BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation
by: Shah, Rutav, et al.
Published: (2024)
by: Shah, Rutav, et al.
Published: (2024)
AirHunt: Bridging VLM Semantics and Continuous Planning for Efficient Aerial Object Navigation
by: Chen, Xuecheng, et al.
Published: (2026)
by: Chen, Xuecheng, et al.
Published: (2026)
Deadlock-Free Hybrid RL-MAPF Framework for Zero-Shot Multi-Robot Navigation
by: Wang, Haoyi, et al.
Published: (2025)
by: Wang, Haoyi, et al.
Published: (2025)
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
by: Du, Zhiying, et al.
Published: (2025)
by: Du, Zhiying, et al.
Published: (2025)
A3VLM: Actionable Articulation-Aware Vision Language Model
by: Huang, Siyuan, et al.
Published: (2024)
by: Huang, Siyuan, et al.
Published: (2024)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Iterative Shaping of Multi-Particle Aggregates based on Action Trees and VLM
by: Lee, Hoi-Yin, et al.
Published: (2025)
by: Lee, Hoi-Yin, et al.
Published: (2025)
Premover: Fast Vision-Language-Action Control by Acting Before Instructions Are Complete
by: Park, Joonha, et al.
Published: (2026)
by: Park, Joonha, et al.
Published: (2026)
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
by: Ma, Teli, et al.
Published: (2026)
by: Ma, Teli, et al.
Published: (2026)
ReplanVLM: Replanning Robotic Tasks with Visual Language Models
by: Mei, Aoran, et al.
Published: (2024)
by: Mei, Aoran, et al.
Published: (2024)
Bridging Deep Reinforcement Learning and Motion Planning for Model-Free Navigation in Cluttered Environments
by: Luo, Licheng, et al.
Published: (2025)
by: Luo, Licheng, et al.
Published: (2025)
SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
by: Zheng, Peiru, et al.
Published: (2025)
by: Zheng, Peiru, et al.
Published: (2025)
MoE-Loco: Mixture of Experts for Multitask Locomotion
by: Huang, Runhan, et al.
Published: (2025)
by: Huang, Runhan, et al.
Published: (2025)
MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation
by: Shen, Yutong, et al.
Published: (2026)
by: Shen, Yutong, et al.
Published: (2026)
Rethinking Intermediate Representation for VLM-based Robot Manipulation
by: Tang, Weiliang, et al.
Published: (2025)
by: Tang, Weiliang, et al.
Published: (2025)
Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models
by: Chen, Canyu, et al.
Published: (2026)
by: Chen, Canyu, et al.
Published: (2026)
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
by: Feng, Qiuxuan, et al.
Published: (2026)
by: Feng, Qiuxuan, et al.
Published: (2026)
Unleashing Humanoid Reaching Potential via Real-world-Ready Skill Space
by: Zhang, Zhikai, et al.
Published: (2025)
by: Zhang, Zhikai, et al.
Published: (2025)
From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
by: Peschl, Markus, et al.
Published: (2025)
by: Peschl, Markus, et al.
Published: (2025)
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
by: Wu, Shijie, et al.
Published: (2024)
by: Wu, Shijie, et al.
Published: (2024)
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
by: Luo, Zekai, et al.
Published: (2025)
by: Luo, Zekai, et al.
Published: (2025)
Unleashing the Potential of Diffusion Models for End-to-End Autonomous Driving
by: Zheng, Yinan, et al.
Published: (2026)
by: Zheng, Yinan, et al.
Published: (2026)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
by: Ma, Teli, et al.
Published: (2025)
by: Ma, Teli, et al.
Published: (2025)
Similar Items
-
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
by: Huang, Zheng, et al.
Published: (2025) -
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025) -
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025) -
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025) -
InterPreT: Interactive Predicate Learning from Language Feedback for Generalizable Task Planning
by: Han, Muzhi, et al.
Published: (2024)