Hume: Introducing System-2 Thinking in Visual-Language-Action Model
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Haoming, Qu, Delin, Yao, Yuanqi, Chen, Qizhi, Lv, Qi, Tang, Yiwen, Shi, Modi, Ren, Guanghui, Yao, Maoqing, Zhao, Bin, Wang, Dong, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025)
by: Yao, Yuanqi, et al.
Published: (2025)
EO-1: An Open Unified Embodied Foundation Model for General Robot Control
by: Qu, Delin, et al.
Published: (2025)
by: Qu, Delin, et al.
Published: (2025)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
by: Bu, Qingwen, et al.
Published: (2025)
by: Bu, Qingwen, et al.
Published: (2025)
Is Diversity All You Need for Scalable Robotic Manipulation?
by: Shi, Modi, et al.
Published: (2025)
by: Shi, Modi, et al.
Published: (2025)
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
by: Yue, Hu, et al.
Published: (2025)
by: Yue, Hu, et al.
Published: (2025)
Adversarial Data Collection: Human-Collaborative Perturbations for Efficient and Robust Robotic Imitation Learning
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
EnerVerse-AC: Envisioning Embodied Environments with Action Condition
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
FreeGaussian: Annotation-free Control of Articulated Objects via 3D Gaussian Splats with Flow Derivatives
by: Chen, Qizhi, et al.
Published: (2024)
by: Chen, Qizhi, et al.
Published: (2024)
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
by: Gao, Xianqiang, et al.
Published: (2026)
by: Gao, Xianqiang, et al.
Published: (2026)
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-System
by: Wei, Yifei, et al.
Published: (2026)
by: Wei, Yifei, et al.
Published: (2026)
Continually Evolving Skill Knowledge in Vision Language Action Model
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation
by: Heng, Liang, et al.
Published: (2025)
by: Heng, Liang, et al.
Published: (2025)
ZeroWBC: Learning Natural Visuomotor Humanoid Control Directly from Human Egocentric Video
by: Yang, Haoran, et al.
Published: (2026)
by: Yang, Haoran, et al.
Published: (2026)
FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset
by: Zhaxizhuoma, et al.
Published: (2024)
by: Zhaxizhuoma, et al.
Published: (2024)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
by: Hu, Yutong, et al.
Published: (2026)
by: Hu, Yutong, et al.
Published: (2026)
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
by: Huang, Zhiyu, et al.
Published: (2026)
by: Huang, Zhiyu, et al.
Published: (2026)
Genie Sim PanoRecon: Fast Immersive Scene Generation from Single-View Panorama
by: Li, Zhijun, et al.
Published: (2026)
by: Li, Zhijun, et al.
Published: (2026)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
by: Wu, Pengyuan, et al.
Published: (2026)
by: Wu, Pengyuan, et al.
Published: (2026)
GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation
by: Qiu, Boxiang, et al.
Published: (2026)
by: Qiu, Boxiang, et al.
Published: (2026)
EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
Trajectory Conditioned Cross-embodiment Skill Transfer
by: Tang, YuHang, et al.
Published: (2025)
by: Tang, YuHang, et al.
Published: (2025)
Concurrent-Allocation Task Execution for Multi-Robot Path-Crossing-Minimal Navigation in Obstacle Environments
by: Hu, Bin-Bin, et al.
Published: (2025)
by: Hu, Bin-Bin, et al.
Published: (2025)
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
by: Li, Ruiying, et al.
Published: (2026)
by: Li, Ruiying, et al.
Published: (2026)
COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models
by: Liu, Kehui, et al.
Published: (2024)
by: Liu, Kehui, et al.
Published: (2024)
Robust Statistics vs. Machine Learning vs. Bayesian Inference: Insights into Handling Faulty GNSS Measurements in Field Robotics
by: Zhang, Haoming
Published: (2025)
by: Zhang, Haoming
Published: (2025)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
Action Contextualization: Adaptive Task Planning and Action Tuning using Large Language Models
by: Gupta, Sthithpragya, et al.
Published: (2024)
by: Gupta, Sthithpragya, et al.
Published: (2024)
Visual Marker Search for Autonomous Drone Landing in Diverse Urban Environments
by: Yao, Jiaohong, et al.
Published: (2026)
by: Yao, Jiaohong, et al.
Published: (2026)
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation
by: Bu, Qingwen, et al.
Published: (2024)
by: Bu, Qingwen, et al.
Published: (2024)
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
by: Liao, Yue, et al.
Published: (2025)
by: Liao, Yue, et al.
Published: (2025)
Interaction Dataset of Autonomous Vehicles with Traffic Lights and Signs
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions
by: Chen, Guanyan, et al.
Published: (2024)
by: Chen, Guanyan, et al.
Published: (2024)
Keypoint Detection Technique for Image-Based Visual Servoing of Manipulators
by: Amiri, Niloufar, et al.
Published: (2024)
by: Amiri, Niloufar, et al.
Published: (2024)
Similar Items
-
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
by: Qu, Delin, et al.
Published: (2025) -
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
by: Yao, Yuanqi, et al.
Published: (2025) -
EO-1: An Open Unified Embodied Foundation Model for General Robot Control
by: Qu, Delin, et al.
Published: (2025) -
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025) -
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)