COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maeda, Koki, Hirasawa, Tosho, Hashimoto, Atsushi, Harashima, Jun, Rybicki, Leszek, Fukasawa, Yusuke, Ushiku, Yoshitaka
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913458541297664
author Maeda, Koki
Hirasawa, Tosho
Hashimoto, Atsushi
Harashima, Jun
Rybicki, Leszek
Fukasawa, Yusuke
Ushiku, Yoshitaka
author_facet Maeda, Koki
Hirasawa, Tosho
Hashimoto, Atsushi
Harashima, Jun
Rybicki, Leszek
Fukasawa, Yusuke
Ushiku, Yoshitaka
contents Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it challenging to query instructional contents from raw video observations. To address this issue, we propose a new dataset, COM Kitchens. The dataset consists of unedited overhead-view videos captured by smartphones, in which participants performed food preparation based on given recipes. Fixed-viewpoint video datasets often lack environmental diversity due to high camera setup costs. We used modern wide-angle smartphone lenses to cover cooking counters from sink to cooktop in an overhead view, capturing activity without in-person assistance. With this setup, we collected a diverse dataset by distributing smartphones to participants. With this dataset, we propose the novel video-to-text retrieval task Online Recipe Retrieval (OnRR) and new video captioning domain Dense Video Captioning on unedited Overhead-View videos (DVC-OV). Our experiments verified the capabilities and limitations of current web-video-based SOTA methods in handling these tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2408_02272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
Maeda, Koki
Hirasawa, Tosho
Hashimoto, Atsushi
Harashima, Jun
Rybicki, Leszek
Fukasawa, Yusuke
Ushiku, Yoshitaka
Computer Vision and Pattern Recognition
Computation and Language
Multimedia
Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it challenging to query instructional contents from raw video observations. To address this issue, we propose a new dataset, COM Kitchens. The dataset consists of unedited overhead-view videos captured by smartphones, in which participants performed food preparation based on given recipes. Fixed-viewpoint video datasets often lack environmental diversity due to high camera setup costs. We used modern wide-angle smartphone lenses to cover cooking counters from sink to cooktop in an overhead view, capturing activity without in-person assistance. With this setup, we collected a diverse dataset by distributing smartphones to participants. With this dataset, we propose the novel video-to-text retrieval task Online Recipe Retrieval (OnRR) and new video captioning domain Dense Video Captioning on unedited Overhead-View videos (DVC-OV). Our experiments verified the capabilities and limitations of current web-video-based SOTA methods in handling these tasks.
title COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
topic Computer Vision and Pattern Recognition
Computation and Language
Multimedia
url https://arxiv.org/abs/2408.02272