MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zang, Yuan, Tan, Hao, Yoon, Seunghyun, Dernoncourt, Franck, Gu, Jiuxiang, Kafle, Kushal, Sun, Chen, Bui, Trung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916794162216960
author Zang, Yuan
Tan, Hao
Yoon, Seunghyun
Dernoncourt, Franck
Gu, Jiuxiang
Kafle, Kushal
Sun, Chen
Bui, Trung
author_facet Zang, Yuan
Tan, Hao
Yoon, Seunghyun
Dernoncourt, Franck
Gu, Jiuxiang
Kafle, Kushal
Sun, Chen
Bui, Trung
contents We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic semantic-level video summarization, and are not suitable for providing step-by-step executable instructions and illustrations, both of which are crucial for instructional videos. We propose a novel benchmark for user interface (UI) instructional video summarization to fill the gap. We collect a dataset of 2,413 UI instructional videos, which spans over 167 hours. These videos are manually annotated for video segmentation, text summarization, and video summarization, which enable the comprehensive evaluations for concise and executable video summarization. We conduct extensive experiments on our collected MS4UI dataset, which suggest that state-of-the-art multi-modal summarization methods struggle on UI video summarization, and highlight the importance of new methods for UI instructional video summarization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
Zang, Yuan
Tan, Hao
Yoon, Seunghyun
Dernoncourt, Franck
Gu, Jiuxiang
Kafle, Kushal
Sun, Chen
Bui, Trung
Computer Vision and Pattern Recognition
Computation and Language
We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic semantic-level video summarization, and are not suitable for providing step-by-step executable instructions and illustrations, both of which are crucial for instructional videos. We propose a novel benchmark for user interface (UI) instructional video summarization to fill the gap. We collect a dataset of 2,413 UI instructional videos, which spans over 167 hours. These videos are manually annotated for video segmentation, text summarization, and video summarization, which enable the comprehensive evaluations for concise and executable video summarization. We conduct extensive experiments on our collected MS4UI dataset, which suggest that state-of-the-art multi-modal summarization methods struggle on UI video summarization, and highlight the importance of new methods for UI instructional video summarization.
title MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.12623