EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Haneji, Yuto, Nishimura, Taichi, Kameko, Hirotaka, Shirai, Keisuke, Yoshida, Tomoya, Kajimura, Keiya, Yamamoto, Koki, Cui, Taiyu, Nishimoto, Tomohiro, Mori, Shinsuke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912511261933568
author Haneji, Yuto
Nishimura, Taichi
Kameko, Hirotaka
Shirai, Keisuke
Yoshida, Tomoya
Kajimura, Keiya
Yamamoto, Koki
Cui, Taiyu
Nishimoto, Tomohiro
Mori, Shinsuke
author_facet Haneji, Yuto
Nishimura, Taichi
Kameko, Hirotaka
Shirai, Keisuke
Yoshida, Tomoya
Kajimura, Keiya
Yamamoto, Koki
Cui, Taiyu
Nishimoto, Tomohiro
Mori, Shinsuke
contents Mistake action detection is crucial for developing intelligent archives that detect workers' errors and provide feedback. Existing studies have focused on visually apparent mistakes in free-style activities, resulting in video-only approaches to mistake detection. However, in text-following activities, models cannot determine the correctness of some actions without referring to the texts. Additionally, current mistake datasets rarely use procedural texts for video recording except for cooking. To fill these gaps, this paper proposes the EgoOops dataset, where egocentric videos record erroneous activities when following procedural texts across diverse domains. It features three types of annotations: video-text alignment, mistake labels, and descriptions for mistakes. We also propose a mistake detection approach, combining video-text alignment and mistake label classification to leverage the texts. Our experimental results show that incorporating procedural texts is essential for mistake detection. Data is available through https://y-haneji.github.io/EgoOops-project-page/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05343
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
Haneji, Yuto
Nishimura, Taichi
Kameko, Hirotaka
Shirai, Keisuke
Yoshida, Tomoya
Kajimura, Keiya
Yamamoto, Koki
Cui, Taiyu
Nishimoto, Tomohiro
Mori, Shinsuke
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Mistake action detection is crucial for developing intelligent archives that detect workers' errors and provide feedback. Existing studies have focused on visually apparent mistakes in free-style activities, resulting in video-only approaches to mistake detection. However, in text-following activities, models cannot determine the correctness of some actions without referring to the texts. Additionally, current mistake datasets rarely use procedural texts for video recording except for cooking. To fill these gaps, this paper proposes the EgoOops dataset, where egocentric videos record erroneous activities when following procedural texts across diverse domains. It features three types of annotations: video-text alignment, mistake labels, and descriptions for mistakes. We also propose a mistake detection approach, combining video-text alignment and mistake label classification to leverage the texts. Our experimental results show that incorporating procedural texts is essential for mistake detection. Data is available through https://y-haneji.github.io/EgoOops-project-page/.
title EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.05343