Aligning Step-by-Step Instructional Diagrams to Video Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiahao, Cherian, Anoop, Liu, Yanbin, Ben-Shabat, Yizhak, Rodriguez, Cristian, Gould, Stephen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916501916745728
author Zhang, Jiahao
Cherian, Anoop
Liu, Yanbin
Ben-Shabat, Yizhak
Rodriguez, Cristian
Gould, Stephen
author_facet Zhang, Jiahao
Cherian, Anoop
Liu, Yanbin
Ben-Shabat, Yizhak
Rodriguez, Cristian
Gould, Stephen
contents Multimodal alignment facilitates the retrieval of instances from one modality when queried using another. In this paper, we consider a novel setting where such an alignment is between (i) instruction steps that are depicted as assembly diagrams (commonly seen in Ikea assembly manuals) and (ii) video segments from in-the-wild videos; these videos comprising an enactment of the assembly actions in the real world. To learn this alignment, we introduce a novel supervised contrastive learning method that learns to align videos with the subtle details in the assembly diagrams, guided by a set of novel losses. To study this problem and demonstrate the effectiveness of our method, we introduce a novel dataset: IAW for Ikea assembly in the wild consisting of 183 hours of videos from diverse furniture assembly collections and nearly 8,300 illustrations from their associated instruction manuals and annotated for their ground truth alignments. We define two tasks on this dataset: First, nearest neighbor retrieval between video segments and illustrations, and, second, alignment of instruction steps and the segments for each video. Extensive experiments on IAW demonstrate superior performances of our approach against alternatives.
format Preprint
id arxiv_https___arxiv_org_abs_2303_13800
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
Zhang, Jiahao
Cherian, Anoop
Liu, Yanbin
Ben-Shabat, Yizhak
Rodriguez, Cristian
Gould, Stephen
Computer Vision and Pattern Recognition
Multimodal alignment facilitates the retrieval of instances from one modality when queried using another. In this paper, we consider a novel setting where such an alignment is between (i) instruction steps that are depicted as assembly diagrams (commonly seen in Ikea assembly manuals) and (ii) video segments from in-the-wild videos; these videos comprising an enactment of the assembly actions in the real world. To learn this alignment, we introduce a novel supervised contrastive learning method that learns to align videos with the subtle details in the assembly diagrams, guided by a set of novel losses. To study this problem and demonstrate the effectiveness of our method, we introduce a novel dataset: IAW for Ikea assembly in the wild consisting of 183 hours of videos from diverse furniture assembly collections and nearly 8,300 illustrations from their associated instruction manuals and annotated for their ground truth alignments. We define two tasks on this dataset: First, nearest neighbor retrieval between video segments and illustrations, and, second, alignment of instruction steps and the segments for each video. Extensive experiments on IAW demonstrate superior performances of our approach against alternatives.
title Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.13800