PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lykov, Artem, Sam, Jeffrin, Nguyen, Hung Khang, Kozlovskiy, Vladislav, Mahmoud, Yara, Serpiva, Valerii, Cabrera, Miguel Altamirano, Konenkov, Mikhail, Tsetserukou, Dzmitry
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909792868499456
author Lykov, Artem
Sam, Jeffrin
Nguyen, Hung Khang
Kozlovskiy, Vladislav
Mahmoud, Yara
Serpiva, Valerii
Cabrera, Miguel Altamirano
Konenkov, Mikhail
Tsetserukou, Dzmitry
author_facet Lykov, Artem
Sam, Jeffrin
Nguyen, Hung Khang
Kozlovskiy, Vladislav
Mahmoud, Yara
Serpiva, Valerii
Cabrera, Miguel Altamirano
Konenkov, Mikhail
Tsetserukou, Dzmitry
contents We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video demonstrations of candidate trajectories, executes them on the robot, and iteratively re-plans in response to failures. This approach enables robust recovery from execution errors. We evaluate PhysicalAgent across multiple perceptual modalities (egocentric, third-person, and simulated) and robotic embodiments (bimanual UR3, Unitree G1 humanoid, simulated GR1), comparing against state-of-the-art task-specific baselines. Experiments demonstrate that our method consistently outperforms prior approaches, achieving up to 83% success on human-familiar tasks. Physical trials reveal that first-attempt success is limited (20-30%), yet iterative correction increases overall success to 80% across platforms. These results highlight the potential of video-based generative reasoning for general-purpose robotic manipulation and underscore the importance of iterative execution for recovering from initial failures. Our framework paves the way for scalable, adaptable, and robust robot control.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13903
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
Lykov, Artem
Sam, Jeffrin
Nguyen, Hung Khang
Kozlovskiy, Vladislav
Mahmoud, Yara
Serpiva, Valerii
Cabrera, Miguel Altamirano
Konenkov, Mikhail
Tsetserukou, Dzmitry
Robotics
We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video demonstrations of candidate trajectories, executes them on the robot, and iteratively re-plans in response to failures. This approach enables robust recovery from execution errors. We evaluate PhysicalAgent across multiple perceptual modalities (egocentric, third-person, and simulated) and robotic embodiments (bimanual UR3, Unitree G1 humanoid, simulated GR1), comparing against state-of-the-art task-specific baselines. Experiments demonstrate that our method consistently outperforms prior approaches, achieving up to 83% success on human-familiar tasks. Physical trials reveal that first-attempt success is limited (20-30%), yet iterative correction increases overall success to 80% across platforms. These results highlight the potential of video-based generative reasoning for general-purpose robotic manipulation and underscore the importance of iterative execution for recovering from initial failures. Our framework paves the way for scalable, adaptable, and robust robot control.
title PhysicalAgent: Towards General Cognitive Robotics with Foundation World Models
topic Robotics
url https://arxiv.org/abs/2509.13903