GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Jiafeng, Jiang, Shixin, Wang, Zekun, Pan, Haojie, Chen, Zerui, Chu, Zheng, Liu, Ming, Fu, Ruiji, Wang, Zhongyuan, Qin, Bing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916301894582272
author Liang, Jiafeng
Jiang, Shixin
Wang, Zekun
Pan, Haojie
Chen, Zerui
Chu, Zheng
Liu, Ming
Fu, Ruiji
Wang, Zhongyuan
Qin, Bing
author_facet Liang, Jiafeng
Jiang, Shixin
Wang, Zekun
Pan, Haojie
Chen, Zerui
Chu, Zheng
Liu, Ming
Fu, Ruiji
Wang, Zhongyuan
Qin, Bing
contents There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at the task level, which can lead to beginners struggling to learn new tasks due to the lack of relevant experience. Moreover, the specific steps without guidelines are trivial and unsystematic, making it difficult to provide a clear tutorial. To address these problems, we present the GUIDE (Guideline-Guided) dataset, which contains 3.5K videos of 560 instructional tasks in 8 domains related to our daily life. Specifically, we annotate each instructional task with a guideline, representing a common pattern shared by all task-related videos. On this basis, we annotate systematic specific steps, including their associated guideline steps, specific step descriptions and timestamps. Our proposed benchmark consists of three sub-tasks to evaluate comprehension ability of models: (1) Step Captioning: models have to generate captions for specific steps from videos. (2) Guideline Summarization: models have to mine the common pattern in task-related videos and summarize a guideline from them. (3) Guideline-Guided Captioning: models have to generate captions for specific steps under the guide of guideline. We evaluate plenty of foundation models with GUIDE and perform in-depth analysis. Given the diversity and practicality of GUIDE, we believe that it can be used as a better benchmark for instructional video comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18227
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
Liang, Jiafeng
Jiang, Shixin
Wang, Zekun
Pan, Haojie
Chen, Zerui
Chu, Zheng
Liu, Ming
Fu, Ruiji
Wang, Zhongyuan
Qin, Bing
Computer Vision and Pattern Recognition
Computation and Language
There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at the task level, which can lead to beginners struggling to learn new tasks due to the lack of relevant experience. Moreover, the specific steps without guidelines are trivial and unsystematic, making it difficult to provide a clear tutorial. To address these problems, we present the GUIDE (Guideline-Guided) dataset, which contains 3.5K videos of 560 instructional tasks in 8 domains related to our daily life. Specifically, we annotate each instructional task with a guideline, representing a common pattern shared by all task-related videos. On this basis, we annotate systematic specific steps, including their associated guideline steps, specific step descriptions and timestamps. Our proposed benchmark consists of three sub-tasks to evaluate comprehension ability of models: (1) Step Captioning: models have to generate captions for specific steps from videos. (2) Guideline Summarization: models have to mine the common pattern in task-related videos and summarize a guideline from them. (3) Guideline-Guided Captioning: models have to generate captions for specific steps under the guide of guideline. We evaluate plenty of foundation models with GUIDE and perform in-depth analysis. Given the diversity and practicality of GUIDE, we believe that it can be used as a better benchmark for instructional video comprehension.
title GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2406.18227