Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
Fuente:
arXiv
Saved in:
| Main Authors: | Lisondra, Matthew, Benhabib, Beno, Nejat, Goldie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
by: Fung, Angus, et al.
Published: (2024)
by: Fung, Angus, et al.
Published: (2024)
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
by: Fung, Angus, et al.
Published: (2022)
by: Fung, Angus, et al.
Published: (2022)
PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living
by: Robinson, Fraser, et al.
Published: (2026)
by: Robinson, Fraser, et al.
Published: (2026)
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024)
by: Ahn, Michael, et al.
Published: (2024)
MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
by: Fung, Angus, et al.
Published: (2024)
by: Fung, Angus, et al.
Published: (2024)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
by: Li, Qixiu, et al.
Published: (2024)
by: Li, Qixiu, et al.
Published: (2024)
LEGENT: Open Platform for Embodied Agents
by: Cheng, Zhili, et al.
Published: (2024)
by: Cheng, Zhili, et al.
Published: (2024)
Real-World Robot Applications of Foundation Models: A Review
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
OceanGym: A Benchmark Environment for Underwater Embodied Agents
by: Xue, Yida, et al.
Published: (2025)
by: Xue, Yida, et al.
Published: (2025)
Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language Use
by: Xi, Jiajun, et al.
Published: (2024)
by: Xi, Jiajun, et al.
Published: (2024)
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
by: Man, Yunze, et al.
Published: (2024)
by: Man, Yunze, et al.
Published: (2024)
EmbodiSwap for Zero-Shot Robot Imitation Learning
by: Dessalene, Eadom, et al.
Published: (2025)
by: Dessalene, Eadom, et al.
Published: (2025)
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
by: Mansour, Malak, et al.
Published: (2025)
by: Mansour, Malak, et al.
Published: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
by: Zheng, Ying, et al.
Published: (2024)
by: Zheng, Ying, et al.
Published: (2024)
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
by: Tavella, Federico, et al.
Published: (2025)
by: Tavella, Federico, et al.
Published: (2025)
ViPRA: Video Prediction for Robot Actions
by: Routray, Sandeep, et al.
Published: (2025)
by: Routray, Sandeep, et al.
Published: (2025)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
by: Shang, Jinghuan, et al.
Published: (2024)
by: Shang, Jinghuan, et al.
Published: (2024)
Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach
by: Tan, Aaron Hao, et al.
Published: (2025)
by: Tan, Aaron Hao, et al.
Published: (2025)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
by: Hu, Yafei, et al.
Published: (2023)
by: Hu, Yafei, et al.
Published: (2023)
Video-Language Critic: Transferable Reward Functions for Language-Conditioned Robotics
by: Alakuijala, Minttu, et al.
Published: (2024)
by: Alakuijala, Minttu, et al.
Published: (2024)
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
EfficientFlow: Efficient Equivariant Flow Policy Learning for Embodied AI
by: Chang, Jianlei, et al.
Published: (2025)
by: Chang, Jianlei, et al.
Published: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
by: Werby, Abdelrhman, et al.
Published: (2024)
by: Werby, Abdelrhman, et al.
Published: (2024)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
by: Blank, Nils, et al.
Published: (2024)
by: Blank, Nils, et al.
Published: (2024)
When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
by: Wu, Tao, et al.
Published: (2025)
by: Wu, Tao, et al.
Published: (2025)
PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
by: Yang, Yandan, et al.
Published: (2024)
by: Yang, Yandan, et al.
Published: (2024)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Composing Pre-Trained Object-Centric Representations for Robotics From "What" and "Where" Foundation Models
by: Shi, Junyao, et al.
Published: (2024)
by: Shi, Junyao, et al.
Published: (2024)
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
by: Wu, Jimmy, et al.
Published: (2024)
by: Wu, Jimmy, et al.
Published: (2024)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
by: Yang, Yue, et al.
Published: (2023)
by: Yang, Yue, et al.
Published: (2023)
Critiques of World Models
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
Similar Items
-
LDTrack: Dynamic People Tracking by Service Robots using Diffusion Models
by: Fung, Angus, et al.
Published: (2024) -
Robots Autonomously Detecting People: A Multimodal Deep Contrastive Learning Method Robust to Intraclass Variations
by: Fung, Angus, et al.
Published: (2022) -
PovNet+: A Deep Learning Architecture for Socially Assistive Robots to Learn and Assist with Multiple Activities of Daily Living
by: Robinson, Fraser, et al.
Published: (2026) -
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
by: Ahn, Michael, et al.
Published: (2024) -
MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models
by: Fung, Angus, et al.
Published: (2024)