LITA: Language Instructed Temporal-Localization Assistant

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, De-An, Liao, Shijia, Radhakrishnan, Subhashree, Yin, Hongxu, Molchanov, Pavlo, Yu, Zhiding, Kautz, Jan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916182238429184
author Huang, De-An
Liao, Shijia
Radhakrishnan, Subhashree
Yin, Hongxu
Molchanov, Pavlo
Yu, Zhiding
Kautz, Jan
author_facet Huang, De-An
Liao, Shijia
Radhakrishnan, Subhashree
Yin, Hongxu
Molchanov, Pavlo
Yu, Zhiding
Kautz, Jan
contents There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the "When?" questions. We identify three key aspects that limit their temporal localization capabilities: (i) time representation, (ii) architecture, and (iii) data. We address these shortcomings by proposing Language Instructed Temporal-Localization Assistant (LITA) with the following features: (1) We introduce time tokens that encode timestamps relative to the video length to better represent time in videos. (2) We introduce SlowFast tokens in the architecture to capture temporal information at fine temporal resolution. (3) We emphasize temporal localization data for LITA. In addition to leveraging existing video datasets with timestamps, we propose a new task, Reasoning Temporal Localization (RTL), along with the dataset, ActivityNet-RTL, for learning and evaluating this task. Reasoning temporal localization requires both the reasoning and temporal localization of Video LLMs. LITA demonstrates strong performance on this challenging task, nearly doubling the temporal mean intersection-over-union (mIoU) of baselines. In addition, we show that our emphasis on temporal localization also substantially improves video-based text generation compared to existing Video LLMs, including a 36% relative improvement of Temporal Understanding. Code is available at: https://github.com/NVlabs/LITA
format Preprint
id arxiv_https___arxiv_org_abs_2403_19046
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LITA: Language Instructed Temporal-Localization Assistant
Huang, De-An
Liao, Shijia
Radhakrishnan, Subhashree
Yin, Hongxu
Molchanov, Pavlo
Yu, Zhiding
Kautz, Jan
Computer Vision and Pattern Recognition
Artificial Intelligence
There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the "When?" questions. We identify three key aspects that limit their temporal localization capabilities: (i) time representation, (ii) architecture, and (iii) data. We address these shortcomings by proposing Language Instructed Temporal-Localization Assistant (LITA) with the following features: (1) We introduce time tokens that encode timestamps relative to the video length to better represent time in videos. (2) We introduce SlowFast tokens in the architecture to capture temporal information at fine temporal resolution. (3) We emphasize temporal localization data for LITA. In addition to leveraging existing video datasets with timestamps, we propose a new task, Reasoning Temporal Localization (RTL), along with the dataset, ActivityNet-RTL, for learning and evaluating this task. Reasoning temporal localization requires both the reasoning and temporal localization of Video LLMs. LITA demonstrates strong performance on this challenging task, nearly doubling the temporal mean intersection-over-union (mIoU) of baselines. In addition, we show that our emphasis on temporal localization also substantially improves video-based text generation compared to existing Video LLMs, including a 36% relative improvement of Temporal Understanding. Code is available at: https://github.com/NVlabs/LITA
title LITA: Language Instructed Temporal-Localization Assistant
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.19046