OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zihao, Cai, Shaofei, Mu, Zhancun, Lin, Haowei, Zhang, Ceyao, Liu, Xuejie, Li, Qing, Liu, Anji, Ma, Xiaojian, Liang, Yitao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Open-World Skill Discovery from Unsegmented Demonstrations
by: Deng, Jingwen, et al.
Published: (2025)
by: Deng, Jingwen, et al.
Published: (2025)
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
MineStudio: A Streamlined Package for Minecraft AI Agent Development
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025)
by: Lian, Kewei, et al.
Published: (2025)
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
by: Li, Muyao, et al.
Published: (2025)
by: Li, Muyao, et al.
Published: (2025)
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Lookahead Path Likelihood Optimization for Diffusion LLMs
by: Liu, Xuejie, et al.
Published: (2026)
by: Liu, Xuejie, et al.
Published: (2026)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
by: Zhao, Guangyu, et al.
Published: (2024)
by: Zhao, Guangyu, et al.
Published: (2024)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Plug-and-Play Context Feature Reuse for Efficient Masked Generation
by: Liu, Xuejie, et al.
Published: (2025)
by: Liu, Xuejie, et al.
Published: (2025)
Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
by: He, Kaichen, et al.
Published: (2025)
by: He, Kaichen, et al.
Published: (2025)
DeFlow: Decoupling Manifold Modeling and Value Maximization for Offline Policy Extraction
by: Mu, Zhancun
Published: (2026)
by: Mu, Zhancun
Published: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
Tractable Transformers for Flexible Conditional Generation
by: Liu, Anji, et al.
Published: (2025)
by: Liu, Anji, et al.
Published: (2025)
The Expressivity Boundary of Probabilistic Circuits: A Comparison with Large Language Models
by: Zhao, Zhiyu, et al.
Published: (2026)
by: Zhao, Zhiyu, et al.
Published: (2026)
UniCode: Augmenting Evaluation for Code Reasoning
by: Zheng, Xinyue, et al.
Published: (2025)
by: Zheng, Xinyue, et al.
Published: (2025)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
A Contextual Combinatorial Bandit Approach to Negotiation
by: Li, Yexin, et al.
Published: (2024)
by: Li, Yexin, et al.
Published: (2024)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2026)
by: Zheng, Minghang, et al.
Published: (2026)
Preserve Support, Not Correspondence: Dynamic Routing for Offline Reinforcement Learning
by: Mu, Zhancun, et al.
Published: (2026)
by: Mu, Zhancun, et al.
Published: (2026)
DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous Grasping
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
OmniSAT: Compact Action Token, Faster Auto Regression
by: Lyu, Huaihai, et al.
Published: (2025)
by: Lyu, Huaihai, et al.
Published: (2025)
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
by: Hirose, Noriaki, et al.
Published: (2025)
by: Hirose, Noriaki, et al.
Published: (2025)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
by: Heo, Miran, et al.
Published: (2025)
by: Heo, Miran, et al.
Published: (2025)
ProAgent: Building Proactive Cooperative Agents with Large Language Models
by: Zhang, Ceyao, et al.
Published: (2023)
by: Zhang, Ceyao, et al.
Published: (2023)
OmniEvent: Unified Event Representation Learning
by: Yan, Weiqi, et al.
Published: (2025)
by: Yan, Weiqi, et al.
Published: (2025)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
by: Glossop, Catherine, et al.
Published: (2025)
by: Glossop, Catherine, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
by: Hu, Xintong, et al.
Published: (2026)
by: Hu, Xintong, et al.
Published: (2026)
JARVIS: A Just-in-Time Augmented Reality VLM-Powered Instruction System for Cross-Reality Task Guidance
by: Sun, Yusi, et al.
Published: (2026)
by: Sun, Yusi, et al.
Published: (2026)
The JARVIS Infrastructure is All You Need for Materials Design
by: Choudhary, Kamal
Published: (2025)
by: Choudhary, Kamal
Published: (2025)
Similar Items
-
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024) -
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025) -
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
by: Cai, Shaofei, et al.
Published: (2024) -
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023) -
Open-World Skill Discovery from Unsegmented Demonstrations
by: Deng, Jingwen, et al.
Published: (2025)