VideoLLM-online: Online Video Large Language Model for Streaming Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Joya, Lv, Zhaoyang, Wu, Shiwei, Lin, Kevin Qinghong, Song, Chenan, Gao, Difei, Liu, Jia-Wei, Gao, Ziteng, Mao, Dongxing, Shou, Mike Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909225565814784
author Chen, Joya
Lv, Zhaoyang
Wu, Shiwei
Lin, Kevin Qinghong
Song, Chenan
Gao, Difei
Liu, Jia-Wei
Gao, Ziteng
Mao, Dongxing
Shou, Mike Zheng
author_facet Chen, Joya
Lv, Zhaoyang
Wu, Shiwei
Lin, Kevin Qinghong
Song, Chenan
Gao, Difei
Liu, Jia-Wei
Gao, Ziteng
Mao, Dongxing
Shou, Mike Zheng
contents Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning methods of these large multimodal models typically treat videos as predetermined clips, making them less effective and efficient at handling streaming video inputs. In this paper, we propose a novel Learning-In-Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time conversation within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training objective designed to perform language modeling for continuous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming dialogue format, and (3) an optimized inference pipeline to speed up the model responses in real-world video streams. With our LIVE framework, we built VideoLLM-online model upon Llama-2/Llama-3 and demonstrate its significant advantages in processing streaming videos. For instance, on average, our model can support streaming dialogue in a 5-minute video clip at over 10 FPS on an A100 GPU. Moreover, it also showcases state-of-the-art performance on public offline video benchmarks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at https://showlab.github.io/videollm-online.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11816
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoLLM-online: Online Video Large Language Model for Streaming Video
Chen, Joya
Lv, Zhaoyang
Wu, Shiwei
Lin, Kevin Qinghong
Song, Chenan
Gao, Difei
Liu, Jia-Wei
Gao, Ziteng
Mao, Dongxing
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning methods of these large multimodal models typically treat videos as predetermined clips, making them less effective and efficient at handling streaming video inputs. In this paper, we propose a novel Learning-In-Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time conversation within a continuous video stream. Our LIVE framework comprises comprehensive approaches to achieve video streaming dialogue, encompassing: (1) a training objective designed to perform language modeling for continuous streaming inputs, (2) a data generation scheme that converts offline temporal annotations into a streaming dialogue format, and (3) an optimized inference pipeline to speed up the model responses in real-world video streams. With our LIVE framework, we built VideoLLM-online model upon Llama-2/Llama-3 and demonstrate its significant advantages in processing streaming videos. For instance, on average, our model can support streaming dialogue in a 5-minute video clip at over 10 FPS on an A100 GPU. Moreover, it also showcases state-of-the-art performance on public offline video benchmarks, such as recognition, captioning, and forecasting. The code, model, data, and demo have been made available at https://showlab.github.io/videollm-online.
title VideoLLM-online: Online Video Large Language Model for Streaming Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11816