A Video-grounded Dialogue Dataset and Metric for Event-driven Activities

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Imrattanatrai, Wiradee, Asada, Masaki, Hasegawa, Kimihiro, Cheng, Zhi-Qi, Fukuda, Ken, Mitamura, Teruko
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916590332674048
author Imrattanatrai, Wiradee
Asada, Masaki
Hasegawa, Kimihiro
Cheng, Zhi-Qi
Fukuda, Ken
Mitamura, Teruko
author_facet Imrattanatrai, Wiradee
Asada, Masaki
Hasegawa, Kimihiro
Cheng, Zhi-Qi
Fukuda, Ken
Mitamura, Teruko
contents This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and more complex video sequences that depict a variety of event-driven activities that require advanced contextual understanding for accurate response generation. The dataset comprises 3,000 dialogues with over 30,000 question-and-answer pairs, derived from 1,000 videos with diverse activity scenarios. VDAct displays a notably challenging characteristic due to its broad spectrum of activity scenarios and wide range of question types. Empirical studies on state-of-the-art vision foundation models highlight their limitations in addressing certain question types on our dataset. Furthermore, VDEval, which integrates dialogue session history and video content summaries extracted from our supplementary Knowledge Graphs to evaluate individual responses, demonstrates a significantly higher correlation with human assessments on the VDAct dataset than existing evaluation metrics that rely solely on the context of single dialogue turns.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
Imrattanatrai, Wiradee
Asada, Masaki
Hasegawa, Kimihiro
Cheng, Zhi-Qi
Fukuda, Ken
Mitamura, Teruko
Computer Vision and Pattern Recognition
Computation and Language
This paper presents VDAct, a dataset for a Video-grounded Dialogue on Event-driven Activities, alongside VDEval, a session-based context evaluation metric specially designed for the task. Unlike existing datasets, VDAct includes longer and more complex video sequences that depict a variety of event-driven activities that require advanced contextual understanding for accurate response generation. The dataset comprises 3,000 dialogues with over 30,000 question-and-answer pairs, derived from 1,000 videos with diverse activity scenarios. VDAct displays a notably challenging characteristic due to its broad spectrum of activity scenarios and wide range of question types. Empirical studies on state-of-the-art vision foundation models highlight their limitations in addressing certain question types on our dataset. Furthermore, VDEval, which integrates dialogue session history and video content summaries extracted from our supplementary Knowledge Graphs to evaluate individual responses, demonstrates a significantly higher correlation with human assessments on the VDAct dataset than existing evaluation metrics that rely solely on the context of single dialogue turns.
title A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2501.18324