Real-World Robot Applications of Foundation Models: A Review
Fuente:
arXiv
Saved in:
| Main Authors: | Kawaharazuka, Kento, Matsushima, Tatsuya, Gambardella, Andrew, Guo, Jiaxian, Paxton, Chris, Zeng, Andy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025)
by: Kawaharazuka, Kento, et al.
Published: (2025)
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024)
by: Kawaharazuka, Kento, et al.
Published: (2024)
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
by: Sendai, Kohei, et al.
Published: (2025)
by: Sendai, Kohei, et al.
Published: (2025)
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
by: Liu, Peiqi, et al.
Published: (2024)
by: Liu, Peiqi, et al.
Published: (2024)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
by: Yang, Yu, et al.
Published: (2025)
by: Yang, Yu, et al.
Published: (2025)
IRASim: A Fine-Grained World Model for Robot Manipulation
by: Zhu, Fangqi, et al.
Published: (2024)
by: Zhu, Fangqi, et al.
Published: (2024)
Real-World Cooking Robot System from Recipes Based on Food State Recognition Using Foundation Models and PDDL
by: Kanazawa, Naoaki, et al.
Published: (2024)
by: Kanazawa, Naoaki, et al.
Published: (2024)
The Safety Challenge of World Models for Embodied AI Agents: A Review
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
by: Huang, Wenlong, et al.
Published: (2026)
by: Huang, Wenlong, et al.
Published: (2026)
Robot Learning from a Physical World Model
by: Mao, Jiageng, et al.
Published: (2025)
by: Mao, Jiageng, et al.
Published: (2025)
MoDem-V2: Visuo-Motor World Models for Real-World Robot Manipulation
by: Lancaster, Patrick, et al.
Published: (2023)
by: Lancaster, Patrick, et al.
Published: (2023)
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
HomeRobot: Open-Vocabulary Mobile Manipulation
by: Yenamandra, Sriram, et al.
Published: (2023)
by: Yenamandra, Sriram, et al.
Published: (2023)
A Systematic Literature Review of Computer Vision Applications in Robotized Wire Harness Assembly
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
General Flow as Foundation Affordance for Scalable Robot Learning
by: Yuan, Chengbo, et al.
Published: (2024)
by: Yuan, Chengbo, et al.
Published: (2024)
EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling
by: Wang, Boyuan, et al.
Published: (2025)
by: Wang, Boyuan, et al.
Published: (2025)
STAR: A Foundation Model-driven Framework for Robust Task Planning and Failure Recovery in Robotic Systems
by: Sakib, Md Sadman, et al.
Published: (2025)
by: Sakib, Md Sadman, et al.
Published: (2025)
World Simulation with Video Foundation Models for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Cosmos World Foundation Model Platform for Physical AI
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
by: Ehsani, Kiana, et al.
Published: (2023)
by: Ehsani, Kiana, et al.
Published: (2023)
Synchronous vs Asynchronous Reinforcement Learning in a Real World Robot
by: Parsaee, Ali, et al.
Published: (2025)
by: Parsaee, Ali, et al.
Published: (2025)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
by: Zhang, Lijun, et al.
Published: (2026)
by: Zhang, Lijun, et al.
Published: (2026)
Rethinking Video Generation Model for the Embodied World
by: Deng, Yufan, et al.
Published: (2026)
by: Deng, Yufan, et al.
Published: (2026)
A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards
by: Patel, Shivansh, et al.
Published: (2025)
by: Patel, Shivansh, et al.
Published: (2025)
Multi-Modal World Model for Physical Robot Interactions: Simultaneous Visual and Tactile Predictions for Enhanced Accuracy
by: Mandil, Willow, et al.
Published: (2023)
by: Mandil, Willow, et al.
Published: (2023)
OPENTOUCH: Bringing Full-Hand Touch to Real-World Interaction
by: Song, Yuxin Ray, et al.
Published: (2025)
by: Song, Yuxin Ray, et al.
Published: (2025)
Conditioning Latent-Space Clusters for Real-World Anomaly Classification
by: Bogdoll, Daniel, et al.
Published: (2023)
by: Bogdoll, Daniel, et al.
Published: (2023)
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
by: Huang, Zilin, et al.
Published: (2026)
by: Huang, Zilin, et al.
Published: (2026)
Shape Completion and Real-Time Visualization in Robotic Ultrasound Spine Acquisitions
by: Gafencu, Miruna-Alexandra, et al.
Published: (2025)
by: Gafencu, Miruna-Alexandra, et al.
Published: (2025)
RealDex: Towards Human-like Grasping for Robotic Dexterous Hand
by: Liu, Yumeng, et al.
Published: (2024)
by: Liu, Yumeng, et al.
Published: (2024)
Helvipad: A Real-World Dataset for Omnidirectional Stereo Depth Estimation
by: Zayene, Mehdi, et al.
Published: (2024)
by: Zayene, Mehdi, et al.
Published: (2024)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
by: Xu, Wenjiang, et al.
Published: (2025)
by: Xu, Wenjiang, et al.
Published: (2025)
Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding
by: Longo, Antonello, et al.
Published: (2025)
by: Longo, Antonello, et al.
Published: (2025)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
by: Shang, Jinghuan, et al.
Published: (2024)
by: Shang, Jinghuan, et al.
Published: (2024)
SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities
by: Jiang, Yunfan, et al.
Published: (2025)
by: Jiang, Yunfan, et al.
Published: (2025)
Chain of World: World Model Thinking in Latent Motion
by: Yang, Fuxiang, et al.
Published: (2026)
by: Yang, Fuxiang, et al.
Published: (2026)
Surfer: Progressive Reasoning with World Models for Robotic Manipulation
by: Ren, Pengzhen, et al.
Published: (2023)
by: Ren, Pengzhen, et al.
Published: (2023)
Similar Items
-
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
by: Kawaharazuka, Kento, et al.
Published: (2025) -
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024) -
Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
by: Kawaharazuka, Kento, et al.
Published: (2024) -
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
by: Sendai, Kohei, et al.
Published: (2025) -
OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
by: Liu, Peiqi, et al.
Published: (2024)