From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
Fuente:
arXiv
Saved in:
| Main Authors: | Szot, Andrew, Mazoure, Bogdan, Attia, Omar, Timofeev, Aleksei, Agrawal, Harsh, Hjelm, Devon, Gan, Zhe, Kira, Zsolt, Toshev, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding Multimodal Large Language Models in Actions
by: Szot, Andrew, et al.
Published: (2024)
by: Szot, Andrew, et al.
Published: (2024)
Large Language Models as Generalizable Policies for Embodied Tasks
by: Szot, Andrew, et al.
Published: (2023)
by: Szot, Andrew, et al.
Published: (2023)
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
by: Sapora, Silvia, et al.
Published: (2025)
by: Sapora, Silvia, et al.
Published: (2025)
Scaling Synthetic Task Generation for Agents via Exploration
by: Ramrakhya, Ram, et al.
Published: (2025)
by: Ramrakhya, Ram, et al.
Published: (2025)
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
by: Klissarov, Martin, et al.
Published: (2024)
by: Klissarov, Martin, et al.
Published: (2024)
Expanding LLM Agent Boundaries with Strategy-Guided Exploration
by: Szot, Andrew, et al.
Published: (2026)
by: Szot, Andrew, et al.
Published: (2026)
The Sandbox Environment for Generalizable Agent Research (SEGAR)
by: Hjelm, R Devon, et al.
Published: (2022)
by: Hjelm, R Devon, et al.
Published: (2022)
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning
by: Ramrakhya, Ram, et al.
Published: (2025)
by: Ramrakhya, Ram, et al.
Published: (2025)
ReLIC: A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI
by: Elawady, Ahmad, et al.
Published: (2024)
by: Elawady, Ahmad, et al.
Published: (2024)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
Reinforcement Learning via Auxiliary Task Distillation
by: Harish, Abhinav Narayan, et al.
Published: (2024)
by: Harish, Abhinav Narayan, et al.
Published: (2024)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
by: Gupta, Gunshi, et al.
Published: (2025)
by: Gupta, Gunshi, et al.
Published: (2025)
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
by: Yadav, Karmesh, et al.
Published: (2025)
by: Yadav, Karmesh, et al.
Published: (2025)
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
SAGE: Sink-Aware Grounded Decoding for Multimodal Hallucination Mitigation
by: Shukla, Tripti, et al.
Published: (2026)
by: Shukla, Tripti, et al.
Published: (2026)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
by: Cheng, Zhili, et al.
Published: (2025)
by: Cheng, Zhili, et al.
Published: (2025)
SIMA 2: A Generalist Embodied Agent for Virtual Worlds
by: SIMA team, et al.
Published: (2025)
by: SIMA team, et al.
Published: (2025)
EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device
by: Chhablani, Gunjan, et al.
Published: (2025)
by: Chhablani, Gunjan, et al.
Published: (2025)
On the benefits of pixel-based hierarchical policies for task generalization
by: Cristea-Platon, Tudor, et al.
Published: (2024)
by: Cristea-Platon, Tudor, et al.
Published: (2024)
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
by: Miao, Kevin, et al.
Published: (2024)
by: Miao, Kevin, et al.
Published: (2024)
Peter L. Berger on Religion
by: Hjelm, Titus
Published: (2026)
by: Hjelm, Titus
Published: (2026)
Uskonto, kieli ja yhteiskunta
by: Hjelm, Titus
Published: (2021)
by: Hjelm, Titus
Published: (2021)
PDB Features
by: Agrawal, Harsh
Published: (2025)
by: Agrawal, Harsh
Published: (2025)
Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation
by: Zhang, Yuhui, et al.
Published: (2023)
by: Zhang, Yuhui, et al.
Published: (2023)
MINTAQA IQTISODIYOTINI BARQAROR RIVOJLANTIRISHDA TURIZM SAMARADORLIGINI OSHIRISHNING XORIJIY TAJRIBALARINING AHAMIYATI
by: Nurbek Toshev
Published: (2025)
by: Nurbek Toshev
Published: (2025)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
OctoNav: Towards Generalist Embodied Navigation
by: Gao, Chen, et al.
Published: (2025)
by: Gao, Chen, et al.
Published: (2025)
An Overview of Recent Developments on Electrodes Modified with Bacteriophages
by: Katarzyna Szot‐Karpińska
Published: (2025)
by: Katarzyna Szot‐Karpińska
Published: (2025)
LA TRANSICIÓN DEMOGRÁFICO-EPIDEMIOLÓGICA EN CHILE, 1960-2001
by: Jorge Szot Meza
Published: (2003)
by: Jorge Szot Meza
Published: (2003)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025)
by: Soni, Aditya Bharat, et al.
Published: (2025)
Specialists or Generalists? Multi-Agent and Single-Agent LLMs for Essay Grading
by: Idowu, Jamiu Adekunle, et al.
Published: (2026)
by: Idowu, Jamiu Adekunle, et al.
Published: (2026)
ClustRecNet: A Novel End-to-End Deep Learning Framework for Clustering Algorithm Recommendation
by: Bakhtyari, Mohammadreza, et al.
Published: (2025)
by: Bakhtyari, Mohammadreza, et al.
Published: (2025)
YOSHLARGA TA'SIR QILUVCHI G'OYAVIY-MAFKURAVIY TAHDIDLARNING YUZAGA KELISHI
by: Toshev, Suhrob Mirzaqulovich
Published: (2025)
by: Toshev, Suhrob Mirzaqulovich
Published: (2025)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Towards Learning a Generalist Model for Embodied Navigation
by: Zheng, Duo, et al.
Published: (2023)
by: Zheng, Duo, et al.
Published: (2023)
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
by: Yang, Longrong, et al.
Published: (2025)
by: Yang, Longrong, et al.
Published: (2025)
Poly-View Contrastive Learning
by: Shidani, Amitis, et al.
Published: (2024)
by: Shidani, Amitis, et al.
Published: (2024)
Embody4D: A Generalist 4D World Model for Embodied AI
by: Tu, Peiyan, et al.
Published: (2026)
by: Tu, Peiyan, et al.
Published: (2026)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
by: Li, Manling, et al.
Published: (2024)
by: Li, Manling, et al.
Published: (2024)
Similar Items
-
Grounding Multimodal Large Language Models in Actions
by: Szot, Andrew, et al.
Published: (2024) -
Large Language Models as Generalizable Policies for Embodied Tasks
by: Szot, Andrew, et al.
Published: (2023) -
GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
by: Sapora, Silvia, et al.
Published: (2025) -
Scaling Synthetic Task Generation for Agents via Exploration
by: Ramrakhya, Ram, et al.
Published: (2025) -
On the Modeling Capabilities of Large Language Models for Sequential Decision Making
by: Klissarov, Martin, et al.
Published: (2024)