GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Shaofei, Zhang, Bowei, Wang, Zihao, Lin, Haowei, Ma, Xiaojian, Liu, Anji, Liang, Yitao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025)
by: Cai, Shaofei, et al.
Published: (2025)
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
MineStudio: A Streamlined Package for Minecraft AI Agent Development
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
Open-World Skill Discovery from Unsegmented Demonstrations
by: Deng, Jingwen, et al.
Published: (2025)
by: Deng, Jingwen, et al.
Published: (2025)
Online Continual Learning For Interactive Instruction Following Agents
by: Kim, Byeonghwi, et al.
Published: (2024)
by: Kim, Byeonghwi, et al.
Published: (2024)
CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
LoopNav: Benchmarking Spatial Consistency in World Models
by: Lian, Kewei, et al.
Published: (2025)
by: Lian, Kewei, et al.
Published: (2025)
Multi-Modal Manipulation via Multi-Modal Policy Consensus
by: Chen, Haonan, et al.
Published: (2025)
by: Chen, Haonan, et al.
Published: (2025)
Cross-Modal Navigation with Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
Model-Based Reinforcement Learning with Multi-Task Offline Pretraining
by: Pan, Minting, et al.
Published: (2023)
by: Pan, Minting, et al.
Published: (2023)
Data Augmentation for Instruction Following Policies via Trajectory Segmentation
by: Höpner, Niklas, et al.
Published: (2025)
by: Höpner, Niklas, et al.
Published: (2025)
Embodied Instruction Following in Unknown Environments
by: Wu, Zhenyu, et al.
Published: (2024)
by: Wu, Zhenyu, et al.
Published: (2024)
LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation
by: Onishchenko, Anatoly O., et al.
Published: (2025)
by: Onishchenko, Anatoly O., et al.
Published: (2025)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
by: Shi, Lucy Xiaoyang, et al.
Published: (2025)
Multi-Level Compositional Reasoning for Interactive Instruction Following
by: Bhambri, Suvaansh, et al.
Published: (2023)
by: Bhambri, Suvaansh, et al.
Published: (2023)
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
by: Zhao, Guangyu, et al.
Published: (2024)
by: Zhao, Guangyu, et al.
Published: (2024)
A Tractable Inference Perspective of Offline RL
by: Liu, Xuejie, et al.
Published: (2023)
by: Liu, Xuejie, et al.
Published: (2023)
Multi-Agent Path Finding in Continuous Spaces with Projected Diffusion Models
by: Liang, Jinhao, et al.
Published: (2024)
by: Liang, Jinhao, et al.
Published: (2024)
Text2Motion: From Natural Language Instructions to Feasible Plans
by: Lin, Kevin, et al.
Published: (2023)
by: Lin, Kevin, et al.
Published: (2023)
STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft
by: Lenzen, Nicholas, et al.
Published: (2024)
by: Lenzen, Nicholas, et al.
Published: (2024)
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents
by: Kim, Byeonghwi, et al.
Published: (2023)
by: Kim, Byeonghwi, et al.
Published: (2023)
Co-jump: Cooperative Jumping with Quadrupedal Robots via Multi-Agent Reinforcement Learning
by: Dong, Shihao, et al.
Published: (2026)
by: Dong, Shihao, et al.
Published: (2026)
MCU: An Evaluation Framework for Open-Ended Game Agents
by: Zheng, Xinyue, et al.
Published: (2023)
by: Zheng, Xinyue, et al.
Published: (2023)
HAIM-DRL: Enhanced Human-in-the-loop Reinforcement Learning for Safe and Efficient Autonomous Driving
by: Huang, Zilin, et al.
Published: (2024)
by: Huang, Zilin, et al.
Published: (2024)
Video-Enhanced Offline Reinforcement Learning: A Model-Based Approach
by: Pan, Minting, et al.
Published: (2025)
by: Pan, Minting, et al.
Published: (2025)
MetaFollower: Adaptable Personalized Autonomous Car Following
by: Chen, Xianda, et al.
Published: (2024)
by: Chen, Xianda, et al.
Published: (2024)
JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction
by: Lin, Fangze, et al.
Published: (2025)
by: Lin, Fangze, et al.
Published: (2025)
On the Exploration of LM-Based Soft Modular Robot Design
by: Ma, Weicheng, et al.
Published: (2024)
by: Ma, Weicheng, et al.
Published: (2024)
AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models
by: Pourkeshavatz, Mozhgan, et al.
Published: (2026)
by: Pourkeshavatz, Mozhgan, et al.
Published: (2026)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
Diffusion-ES: Gradient-free Planning with Diffusion for Autonomous Driving and Zero-Shot Instruction Following
by: Yang, Brian, et al.
Published: (2024)
by: Yang, Brian, et al.
Published: (2024)
ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
by: Hu, Zichao, et al.
Published: (2025)
by: Hu, Zichao, et al.
Published: (2025)
Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
MRS: Multi-Resolution Skills for HRL Agents
by: Sharma, Shashank, et al.
Published: (2025)
by: Sharma, Shashank, et al.
Published: (2025)
Boundary-to-Region Supervision for Offline Safe Reinforcement Learning
by: Su, Huikang, et al.
Published: (2025)
by: Su, Huikang, et al.
Published: (2025)
Multi-Robot Path Planning Combining Heuristics and Multi-Agent Reinforcement Learning
by: Peng, Shaoming
Published: (2023)
by: Peng, Shaoming
Published: (2023)
Similar Items
-
Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
by: Cai, Shaofei, et al.
Published: (2025) -
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
by: Wang, Zihao, et al.
Published: (2024) -
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
by: Cai, Shaofei, et al.
Published: (2025) -
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023) -
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
by: Wang, Zihao, et al.
Published: (2024)