VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Shicheng, Li, Lei, Ren, Shuhuai, Liu, Yuanxin, Liu, Yi, Gao, Rundong, Sun, Xu, Hou, Lu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917780754792448
author Li, Shicheng
Li, Lei
Ren, Shuhuai
Liu, Yuanxin
Liu, Yi
Gao, Rundong
Sun, Xu
Hou, Lu
author_facet Li, Shicheng
Li, Lei
Ren, Shuhuai
Liu, Yuanxin
Liu, Yi
Gao, Rundong
Sun, Xu
Hou, Lu
contents The ability to perceive how objects change over time is a crucial ingredient in human intelligence. However, current benchmarks cannot faithfully reflect the temporal understanding abilities of video-language models (VidLMs) due to the existence of static visual shortcuts. To remedy this issue, we present VITATECS, a diagnostic VIdeo-Text dAtaset for the evaluation of TEmporal Concept underStanding. Specifically, we first introduce a fine-grained taxonomy of temporal concepts in natural language in order to diagnose the capability of VidLMs to comprehend different temporal aspects. Furthermore, to disentangle the correlation between static and temporal information, we generate counterfactual video descriptions that differ from the original one only in the specified temporal aspect. We employ a semi-automatic data collection framework using large language models and human-in-the-loop annotation to obtain high-quality counterfactual descriptions efficiently. Evaluation of representative video-language understanding models confirms their deficiency in temporal understanding, revealing the need for greater emphasis on the temporal elements in video-language research.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17404
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
Li, Shicheng
Li, Lei
Ren, Shuhuai
Liu, Yuanxin
Liu, Yi
Gao, Rundong
Sun, Xu
Hou, Lu
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
The ability to perceive how objects change over time is a crucial ingredient in human intelligence. However, current benchmarks cannot faithfully reflect the temporal understanding abilities of video-language models (VidLMs) due to the existence of static visual shortcuts. To remedy this issue, we present VITATECS, a diagnostic VIdeo-Text dAtaset for the evaluation of TEmporal Concept underStanding. Specifically, we first introduce a fine-grained taxonomy of temporal concepts in natural language in order to diagnose the capability of VidLMs to comprehend different temporal aspects. Furthermore, to disentangle the correlation between static and temporal information, we generate counterfactual video descriptions that differ from the original one only in the specified temporal aspect. We employ a semi-automatic data collection framework using large language models and human-in-the-loop annotation to obtain high-quality counterfactual descriptions efficiently. Evaluation of representative video-language understanding models confirms their deficiency in temporal understanding, revealing the need for greater emphasis on the temporal elements in video-language research.
title VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2311.17404