FAST: Efficient Action Tokenization for Vision-Language-Action Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pertsch, Karl, Stachowicz, Kyle, Ichter, Brian, Driess, Danny, Nair, Suraj, Vuong, Quan, Mees, Oier, Finn, Chelsea, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
von: Torne, Marcel, et al.
Veröffentlicht: (2026)
von: Torne, Marcel, et al.
Veröffentlicht: (2026)
Training Strategies for Efficient Embodied Reasoning
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Emergence of Human to Robot Transfer in Vision-Language-Action Models
von: Kareer, Simar, et al.
Veröffentlicht: (2025)
von: Kareer, Simar, et al.
Veröffentlicht: (2025)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
von: Gao, Tian, et al.
Veröffentlicht: (2026)
von: Gao, Tian, et al.
Veröffentlicht: (2026)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2025)
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2025)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
von: Driess, Danny, et al.
Veröffentlicht: (2025)
von: Driess, Danny, et al.
Veröffentlicht: (2025)
Robotic Control via Embodied Chain-of-Thought Reasoning
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
von: Zawalski, Michał, et al.
Veröffentlicht: (2024)
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
von: Black, Kevin, et al.
Veröffentlicht: (2024)
von: Black, Kevin, et al.
Veröffentlicht: (2024)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
von: Chen, William, et al.
Veröffentlicht: (2026)
von: Chen, William, et al.
Veröffentlicht: (2026)
Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
von: Jones, Joshua, et al.
Veröffentlicht: (2025)
von: Jones, Joshua, et al.
Veröffentlicht: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2024)
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
von: Lee, Tony, et al.
Veröffentlicht: (2026)
von: Lee, Tony, et al.
Veröffentlicht: (2026)
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
von: Intelligence, Physical, et al.
Veröffentlicht: (2025)
von: Intelligence, Physical, et al.
Veröffentlicht: (2025)
Robust Finetuning of Vision-Language-Action Robot Policies via Parameter Merging
von: Yadav, Yajat, et al.
Veröffentlicht: (2025)
von: Yadav, Yajat, et al.
Veröffentlicht: (2025)
RACER: Epistemic Risk-Sensitive RL Enables Fast Driving with Fewer Crashes
von: Stachowicz, Kyle, et al.
Veröffentlicht: (2024)
von: Stachowicz, Kyle, et al.
Veröffentlicht: (2024)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
von: Nakamoto, Mitsuhiko, et al.
Veröffentlicht: (2024)
von: Nakamoto, Mitsuhiko, et al.
Veröffentlicht: (2024)
Evaluating Real-World Robot Manipulation Policies in Simulation
von: Li, Xuanlin, et al.
Veröffentlicht: (2024)
von: Li, Xuanlin, et al.
Veröffentlicht: (2024)
Policy Adaptation via Language Optimization: Decomposing Tasks for Few-Shot Imitation
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos
von: Hirose, Noriaki, et al.
Veröffentlicht: (2024)
von: Hirose, Noriaki, et al.
Veröffentlicht: (2024)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
von: Doshi, Ria, et al.
Veröffentlicht: (2024)
von: Doshi, Ria, et al.
Veröffentlicht: (2024)
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
von: Dong, Perry, et al.
Veröffentlicht: (2026)
von: Dong, Perry, et al.
Veröffentlicht: (2026)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
von: Xu, Charles, et al.
Veröffentlicht: (2026)
von: Xu, Charles, et al.
Veröffentlicht: (2026)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
von: Kim, Moo Jin, et al.
Veröffentlicht: (2025)
Yell At Your Robot: Improving On-the-Fly from Language Corrections
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2024)
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2024)
Octo: An Open-Source Generalist Robot Policy
von: Octo Model Team, et al.
Veröffentlicht: (2024)
von: Octo Model Team, et al.
Veröffentlicht: (2024)
Self-Guided Action Diffusion
von: Malhotra, Rhea, et al.
Veröffentlicht: (2025)
von: Malhotra, Rhea, et al.
Veröffentlicht: (2025)
VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model
von: Guo, Yanjiang, et al.
Veröffentlicht: (2026)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2026)
Autonomous Improvement of Instruction Following Skills via Foundation Models
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2024)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
von: Glossop, Catherine, et al.
Veröffentlicht: (2025)
von: Glossop, Catherine, et al.
Veröffentlicht: (2025)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Action Tokenizer Matters in In-Context Imitation Learning
von: Vuong, An Dinh, et al.
Veröffentlicht: (2025)
von: Vuong, An Dinh, et al.
Veröffentlicht: (2025)
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
von: Hirose, Noriaki, et al.
Veröffentlicht: (2025)
von: Hirose, Noriaki, et al.
Veröffentlicht: (2025)
GHIL-Glue: Hierarchical Control with Filtered Subgoal Images
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
von: Hatch, Kyle B., et al.
Veröffentlicht: (2024)
The Ingredients for Robotic Diffusion Transformers
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
von: Dasari, Sudeep, et al.
Veröffentlicht: (2024)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
von: Zhong, Yifan, et al.
Veröffentlicht: (2025)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
von: Chen, William, et al.
Veröffentlicht: (2024)
von: Chen, William, et al.
Veröffentlicht: (2024)
Pushing the Limits of Cross-Embodiment Learning for Manipulation and Navigation
von: Yang, Jonathan, et al.
Veröffentlicht: (2024)
von: Yang, Jonathan, et al.
Veröffentlicht: (2024)
Commonsense Reasoning for Legged Robot Adaptation with Vision-Language Models
von: Chen, Annie S., et al.
Veröffentlicht: (2024)
von: Chen, Annie S., et al.
Veröffentlicht: (2024)
mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
von: Pai, Jonas, et al.
Veröffentlicht: (2025)
ALOHA Unleashed: A Simple Recipe for Robot Dexterity
von: Zhao, Tony Z., et al.
Veröffentlicht: (2024)
von: Zhao, Tony Z., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
von: Torne, Marcel, et al.
Veröffentlicht: (2026) -
Training Strategies for Efficient Embodied Reasoning
von: Chen, William, et al.
Veröffentlicht: (2025) -
Emergence of Human to Robot Transfer in Vision-Language-Action Models
von: Kareer, Simar, et al.
Veröffentlicht: (2025) -
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
von: Gao, Tian, et al.
Veröffentlicht: (2026) -
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2025)