Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yang, Wang, Chenwei, Lu, Ouyang, Zhao, Yuan, Ge, Yunfei, Sun, Zhenglong, Li, Xiu, Zhang, Chi, Bai, Chenjia, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025)
by: Yang, Siyuan, et al.
Published: (2025)
Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning
by: Xie, Weiji, et al.
Published: (2025)
by: Xie, Weiji, et al.
Published: (2025)
Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
by: Yin, Shiyuan, et al.
Published: (2025)
by: Yin, Shiyuan, et al.
Published: (2025)
Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning
by: Shi, Jiyuan, et al.
Published: (2025)
by: Shi, Jiyuan, et al.
Published: (2025)
Learning Soccer Skills for Humanoid Robots: A Progressive Perception-Action Framework
by: Kong, Jipeng, et al.
Published: (2026)
by: Kong, Jipeng, et al.
Published: (2026)
Preference Aligned Diffusion Planner for Quadrupedal Locomotion Control
by: Yuan, Xinyi, et al.
Published: (2024)
by: Yuan, Xinyi, et al.
Published: (2024)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
DeCoNav: Dialog enhanced Long-Horizon Collaborative Vision-Language Navigation
by: Zhou, Sunyao, et al.
Published: (2026)
by: Zhou, Sunyao, et al.
Published: (2026)
VLP: Vision-Language Preference Learning for Embodied Manipulation
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
by: Fan, Chenyou, et al.
Published: (2025)
by: Fan, Chenyou, et al.
Published: (2025)
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation
by: Zhang, Junjie, et al.
Published: (2024)
by: Zhang, Junjie, et al.
Published: (2024)
X-Loco: Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation
by: Wang, Dewei, et al.
Published: (2026)
by: Wang, Dewei, et al.
Published: (2026)
HUSKY: Humanoid Skateboarding System via Physics-Aware Whole-Body Control
by: Han, Jinrui, et al.
Published: (2026)
by: Han, Jinrui, et al.
Published: (2026)
OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
by: Cheng, Zhengxue, et al.
Published: (2025)
by: Cheng, Zhengxue, et al.
Published: (2025)
PCHC: Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning
by: Li, Huanyu, et al.
Published: (2026)
by: Li, Huanyu, et al.
Published: (2026)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
by: Luo, Hao, et al.
Published: (2026)
by: Luo, Hao, et al.
Published: (2026)
TextOp: Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control
by: Xie, Weiji, et al.
Published: (2026)
by: Xie, Weiji, et al.
Published: (2026)
Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-Training
by: He, Haoran, et al.
Published: (2024)
by: He, Haoran, et al.
Published: (2024)
Pro-HOI: Perceptive Root-guided Humanoid-Object Interaction
by: Lin, Yuhang, et al.
Published: (2026)
by: Lin, Yuhang, et al.
Published: (2026)
Beyond Short-Horizon: VQ-Memory for Robust Long-Horizon Manipulation in Non-Markovian Simulation Benchmarks
by: Wang, Honghui, et al.
Published: (2026)
by: Wang, Honghui, et al.
Published: (2026)
MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains
by: Wang, Dewei, et al.
Published: (2025)
by: Wang, Dewei, et al.
Published: (2025)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
by: Li, Zuolei, et al.
Published: (2025)
by: Li, Zuolei, et al.
Published: (2025)
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
by: Xie, Weiji, et al.
Published: (2025)
by: Xie, Weiji, et al.
Published: (2025)
Skill-Nav: Enhanced Navigation with Versatile Quadrupedal Locomotion via Waypoint Interface
by: Wang, Dewei, et al.
Published: (2025)
by: Wang, Dewei, et al.
Published: (2025)
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
by: Zheng, Yupeng, et al.
Published: (2026)
by: Zheng, Yupeng, et al.
Published: (2026)
RynnVLA-002: A Unified Vision-Language-Action and World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
by: Liu, Yudong, et al.
Published: (2026)
by: Liu, Yudong, et al.
Published: (2026)
RotVLA: Rotational Latent Action for Vision-Language-Action Model
by: Li, Qiwei, et al.
Published: (2026)
by: Li, Qiwei, et al.
Published: (2026)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
by: Chen, Xiaoyu, et al.
Published: (2025)
by: Chen, Xiaoyu, et al.
Published: (2025)
ACG: Action Coherence Guidance for Flow-based Vision-Language-Action models
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
by: Zhu, Xiang, et al.
Published: (2026)
by: Zhu, Xiang, et al.
Published: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
Multiagent Reinforcement Learning with Neighbor Action Estimation
by: Luo, Zhenglong, et al.
Published: (2026)
by: Luo, Zhenglong, et al.
Published: (2026)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
by: Fan, Shichao, et al.
Published: (2025)
by: Fan, Shichao, et al.
Published: (2025)
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
Similar Items
-
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
by: Yang, Siyuan, et al.
Published: (2025) -
Humanoid Whole-Body Locomotion on Narrow Terrain via Dynamic Balance and Reinforcement Learning
by: Xie, Weiji, et al.
Published: (2025) -
Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
by: Yin, Shiyuan, et al.
Published: (2025) -
Adversarial Locomotion and Motion Imitation for Humanoid Policy Learning
by: Shi, Jiyuan, et al.
Published: (2025) -
Learning Soccer Skills for Humanoid Robots: A Progressive Perception-Action Framework
by: Kong, Jipeng, et al.
Published: (2026)