Multimedia Generative Script Learning for Task Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qingyun, Li, Manling, Chan, Hou Pong, Huang, Lifu, Hockenmaier, Julia, Chowdhary, Girish, Ji, Heng
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915333988679680
author Wang, Qingyun
Li, Manling
Chan, Hou Pong
Huang, Lifu
Hockenmaier, Julia
Chowdhary, Girish
Ji, Heng
author_facet Wang, Qingyun
Li, Manling
Chan, Hou Pong
Huang, Lifu
Hockenmaier, Julia
Chowdhary, Girish
Ji, Heng
contents Goal-oriented generative script learning aims to generate subsequent steps to reach a particular goal, which is an essential task to assist robots or humans in performing stereotypical activities. An important aspect of this process is the ability to capture historical states visually, which provides detailed information that is not covered by text and will guide subsequent steps. Therefore, we propose a new task, Multimedia Generative Script Learning, to generate subsequent steps by tracking historical states in both text and vision modalities, as well as presenting the first benchmark containing 5,652 tasks and 79,089 multimedia steps. This task is challenging in three aspects: the multimedia challenge of capturing the visual states in images, the induction challenge of performing unseen tasks, and the diversity challenge of covering different information in individual steps. We propose to encode visual state changes through a selective multimedia encoder to address the multimedia challenge, transfer knowledge from previously observed tasks using a retrieval-augmented decoder to overcome the induction challenge, and further present distinct information at each step by optimizing a diversity-oriented contrastive learning objective. We define metrics to evaluate both generation and inductive quality. Experiment results demonstrate that our approach significantly outperforms strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2208_12306
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Multimedia Generative Script Learning for Task Planning
Wang, Qingyun
Li, Manling
Chan, Hou Pong
Huang, Lifu
Hockenmaier, Julia
Chowdhary, Girish
Ji, Heng
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Goal-oriented generative script learning aims to generate subsequent steps to reach a particular goal, which is an essential task to assist robots or humans in performing stereotypical activities. An important aspect of this process is the ability to capture historical states visually, which provides detailed information that is not covered by text and will guide subsequent steps. Therefore, we propose a new task, Multimedia Generative Script Learning, to generate subsequent steps by tracking historical states in both text and vision modalities, as well as presenting the first benchmark containing 5,652 tasks and 79,089 multimedia steps. This task is challenging in three aspects: the multimedia challenge of capturing the visual states in images, the induction challenge of performing unseen tasks, and the diversity challenge of covering different information in individual steps. We propose to encode visual state changes through a selective multimedia encoder to address the multimedia challenge, transfer knowledge from previously observed tasks using a retrieval-augmented decoder to overcome the induction challenge, and further present distinct information at each step by optimizing a diversity-oriented contrastive learning objective. We define metrics to evaluate both generation and inductive quality. Experiment results demonstrate that our approach significantly outperforms strong baselines.
title Multimedia Generative Script Learning for Task Planning
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2208.12306