Chrono: A Simple Blueprint for Representing Time in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rodriguez, Hector, Meinardus, Boris, Batra, Anil, Rohrbach, Anna, Rohrbach, Marcus
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909978289242112
author Rodriguez, Hector
Meinardus, Boris
Batra, Anil
Rohrbach, Anna
Rohrbach, Marcus
author_facet Rodriguez, Hector
Meinardus, Boris
Batra, Anil
Rohrbach, Anna
Rohrbach, Marcus
contents The recent success of Large Language Models (LLMs) has prompted the extension to the multimodal domain, developing image-text Multimodal LLMs (MLLMs) and then video-text models. In this work, we investigate the challenge of contextual and temporal comprehension in video-language models by exploring the task of temporal localization in videos. To address this problem, prior works have developed complex task-specific architectures, novel modules to embed time into MLLMs, or leveraged additional input signals such as video transcripts to best encode contextual and temporal information. We find that most of these efforts are surpassed by a much simpler design. We introduce Chrono, a universal sequence blueprint that can be applied to any image-text pretrained MLLM. In extensive experiments spanning different MLLM architectures and sizes, finetuning and zero-shot settings, we demonstrate new state-of-the-art results in moment retrieval on the widely used benchmarks Charades-STA, QVHighlights, and ActivityNet Captions, as well as in grounded video question answering on NExT-GQA.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chrono: A Simple Blueprint for Representing Time in MLLMs
Rodriguez, Hector
Meinardus, Boris
Batra, Anil
Rohrbach, Anna
Rohrbach, Marcus
Computer Vision and Pattern Recognition
The recent success of Large Language Models (LLMs) has prompted the extension to the multimodal domain, developing image-text Multimodal LLMs (MLLMs) and then video-text models. In this work, we investigate the challenge of contextual and temporal comprehension in video-language models by exploring the task of temporal localization in videos. To address this problem, prior works have developed complex task-specific architectures, novel modules to embed time into MLLMs, or leveraged additional input signals such as video transcripts to best encode contextual and temporal information. We find that most of these efforts are surpassed by a much simpler design. We introduce Chrono, a universal sequence blueprint that can be applied to any image-text pretrained MLLM. In extensive experiments spanning different MLLM architectures and sizes, finetuning and zero-shot settings, we demonstrate new state-of-the-art results in moment retrieval on the widely used benchmarks Charades-STA, QVHighlights, and ActivityNet Captions, as well as in grounded video question answering on NExT-GQA.
title Chrono: A Simple Blueprint for Representing Time in MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.18113