ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hasegawa, Kimihiro, Imrattanatrai, Wiradee, Cheng, Zhi-Qi, Asada, Masaki, Holm, Susan, Wang, Yuran, Fukuda, Ken, Mitamura, Teruko
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915595868438528
author Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Cheng, Zhi-Qi
Asada, Masaki
Holm, Susan
Wang, Yuran
Fukuda, Ken
Mitamura, Teruko
author_facet Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Cheng, Zhi-Qi
Asada, Masaki
Holm, Susan
Wang, Yuran
Fukuda, Ken
Mitamura, Teruko
contents Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification tasks, e.g., action recognition or temporal action segmentation. In this paper, we present a novel evaluation dataset, ProMQA, to measure system advancements in application-oriented scenarios. ProMQA consists of 401 multimodal procedural QA pairs on user recording of procedural activities, i.e., cooking, coupled with their corresponding instructions/recipes. For QA annotation, we take a cost-effective human-LLM collaborative approach, where the existing annotation is augmented with LLM-generated QA pairs that are later verified by humans. We then provide the benchmark results to set the baseline performance on ProMQA. Our experiment reveals a significant gap between human performance and that of current systems, including competitive proprietary multimodal models. We hope our dataset sheds light on new aspects of models' multimodal understanding capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22211
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Cheng, Zhi-Qi
Asada, Masaki
Holm, Susan
Wang, Yuran
Fukuda, Ken
Mitamura, Teruko
Computation and Language
Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification tasks, e.g., action recognition or temporal action segmentation. In this paper, we present a novel evaluation dataset, ProMQA, to measure system advancements in application-oriented scenarios. ProMQA consists of 401 multimodal procedural QA pairs on user recording of procedural activities, i.e., cooking, coupled with their corresponding instructions/recipes. For QA annotation, we take a cost-effective human-LLM collaborative approach, where the existing annotation is augmented with LLM-generated QA pairs that are later verified by humans. We then provide the benchmark results to set the baseline performance on ProMQA. Our experiment reveals a significant gap between human performance and that of current systems, including competitive proprietary multimodal models. We hope our dataset sheds light on new aspects of models' multimodal understanding capabilities.
title ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding
topic Computation and Language
url https://arxiv.org/abs/2410.22211