Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Salimpour, Sahar, Fu, Lei, Rachwał, Kajetan, Bertrand, Pascal, O'Sullivan, Kevin, Jakob, Robert, Keramat, Farhad, Militano, Leonardo, Toffetti, Giovanni, Edelman, Harry, Queralta, Jorge Peña
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914153865674752
author Salimpour, Sahar
Fu, Lei
Rachwał, Kajetan
Bertrand, Pascal
O'Sullivan, Kevin
Jakob, Robert
Keramat, Farhad
Militano, Leonardo
Toffetti, Giovanni
Edelman, Harry
Queralta, Jorge Peña
author_facet Salimpour, Sahar
Fu, Lei
Rachwał, Kajetan
Bertrand, Pascal
O'Sullivan, Kevin
Jakob, Robert
Keramat, Farhad
Militano, Leonardo
Toffetti, Giovanni
Edelman, Harry
Queralta, Jorge Peña
contents Foundation models, including large language models (LLMs) and vision-language models (VLMs), have recently enabled novel approaches to robot autonomy and human-robot interfaces. In parallel, vision-language-action models (VLAs) or large behavior models (LBMs) are increasing the dexterity and capabilities of robotic systems. This survey paper reviews works that advance agentic applications and architectures, including initial efforts with GPT-style interfaces and more complex systems where AI agents function as coordinators, planners, perception actors, or generalist interfaces. Such agentic architectures allow robots to reason over natural language instructions, invoke APIs, plan task sequences, or assist in operations and diagnostics. In addition to peer-reviewed research, due to the fast-evolving nature of the field, we highlight and include community-driven projects, ROS packages, and industrial frameworks that show emerging trends. We propose a taxonomy for classifying model integration approaches and present a comparative analysis of the role that agents play in different solutions in today's literature.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
Salimpour, Sahar
Fu, Lei
Rachwał, Kajetan
Bertrand, Pascal
O'Sullivan, Kevin
Jakob, Robert
Keramat, Farhad
Militano, Leonardo
Toffetti, Giovanni
Edelman, Harry
Queralta, Jorge Peña
Robotics
Artificial Intelligence
Machine Learning
Foundation models, including large language models (LLMs) and vision-language models (VLMs), have recently enabled novel approaches to robot autonomy and human-robot interfaces. In parallel, vision-language-action models (VLAs) or large behavior models (LBMs) are increasing the dexterity and capabilities of robotic systems. This survey paper reviews works that advance agentic applications and architectures, including initial efforts with GPT-style interfaces and more complex systems where AI agents function as coordinators, planners, perception actors, or generalist interfaces. Such agentic architectures allow robots to reason over natural language instructions, invoke APIs, plan task sequences, or assist in operations and diagnostics. In addition to peer-reviewed research, due to the fast-evolving nature of the field, we highlight and include community-driven projects, ROS packages, and industrial frameworks that show emerging trends. We propose a taxonomy for classifying model integration approaches and present a comparative analysis of the role that agents play in different solutions in today's literature.
title Towards Embodied Agentic AI: Review and Classification of LLM- and VLM-Driven Robot Autonomy and Interaction
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.05294