On the Consistency of Video Large Language Models in Temporal Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jung, Minjoon, Xiao, Junbin, Zhang, Byoung-Tak, Yao, Angela
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909538179874816
author Jung, Minjoon
Xiao, Junbin
Zhang, Byoung-Tak
Yao, Angela
author_facet Jung, Minjoon
Xiao, Junbin
Zhang, Byoung-Tak
Yao, Angela
contents Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of temporal grounding. After the model identifies an initial moment within the video content, we apply a series of probes to check if the model's responses align with this initial grounding as an indicator of reliable comprehension. Our results reveal that current Video-LLMs are sensitive to variations in video contents, language queries, and task settings, unveiling severe deficiencies in maintaining consistency. We further explore common prompting and instruction-tuning methods as potential solutions, but find that their improvements are often unstable. To that end, we propose event temporal verification tuning that explicitly accounts for consistency, and demonstrate significant improvements for both grounding and consistency. Our data and code are open-sourced at https://github.com/minjoong507/Consistency-of-Video-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12951
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Consistency of Video Large Language Models in Temporal Comprehension
Jung, Minjoon
Xiao, Junbin
Zhang, Byoung-Tak
Yao, Angela
Computer Vision and Pattern Recognition
Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction consistency -- a key indicator for robustness and trustworthiness of temporal grounding. After the model identifies an initial moment within the video content, we apply a series of probes to check if the model's responses align with this initial grounding as an indicator of reliable comprehension. Our results reveal that current Video-LLMs are sensitive to variations in video contents, language queries, and task settings, unveiling severe deficiencies in maintaining consistency. We further explore common prompting and instruction-tuning methods as potential solutions, but find that their improvements are often unstable. To that end, we propose event temporal verification tuning that explicitly accounts for consistency, and demonstrate significant improvements for both grounding and consistency. Our data and code are open-sourced at https://github.com/minjoong507/Consistency-of-Video-LLM.
title On the Consistency of Video Large Language Models in Temporal Comprehension
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.12951