Egocentric Vision Language Planning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fang, Zhirui, Yang, Ming, Zeng, Weishuai, Li, Boyu, Yue, Junpeng, Ding, Ziluo, Li, Xiu, Lu, Zongqing
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917746288099328
author Fang, Zhirui
Yang, Ming
Zeng, Weishuai
Li, Boyu
Yue, Junpeng
Ding, Ziluo
Li, Xiu
Lu, Zongqing
author_facet Fang, Zhirui
Yang, Ming
Zeng, Weishuai
Li, Boyu
Yue, Junpeng
Ding, Ziluo
Li, Xiu
Lu, Zongqing
contents We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstractions but struggle with grounding in the physical world, often failing to accurately identify object positions in images. A bridge is needed to connect LMMs to the physical world. The paper proposes a novel approach, egocentric vision language planning (EgoPlan), to handle long-horizon tasks from an egocentric perspective in varying household scenarios. This model leverages a diffusion model to simulate the fundamental dynamics between states and actions, integrating techniques like style transfer and optical flow to enhance generalization across different environmental dynamics. The LMM serves as a planner, breaking down instructions into sub-goals and selecting actions based on their alignment with these sub-goals, thus enabling more generalized and effective decision-making. Experiments show that EgoPlan improves long-horizon task success rates from the egocentric view compared to baselines across household scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2408_05802
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Egocentric Vision Language Planning
Fang, Zhirui
Yang, Ming
Zeng, Weishuai
Li, Boyu
Yue, Junpeng
Ding, Ziluo
Li, Xiu
Lu, Zongqing
Computer Vision and Pattern Recognition
We explore leveraging large multi-modal models (LMMs) and text2image models to build a more general embodied agent. LMMs excel in planning long-horizon tasks over symbolic abstractions but struggle with grounding in the physical world, often failing to accurately identify object positions in images. A bridge is needed to connect LMMs to the physical world. The paper proposes a novel approach, egocentric vision language planning (EgoPlan), to handle long-horizon tasks from an egocentric perspective in varying household scenarios. This model leverages a diffusion model to simulate the fundamental dynamics between states and actions, integrating techniques like style transfer and optical flow to enhance generalization across different environmental dynamics. The LMM serves as a planner, breaking down instructions into sub-goals and selecting actions based on their alignment with these sub-goals, thus enabling more generalized and effective decision-making. Experiments show that EgoPlan improves long-horizon task success rates from the egocentric view compared to baselines across household scenarios.
title Egocentric Vision Language Planning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.05802