Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Jiajun, Wang, Yibing, Ma, Hanghang, Wu, Xiaoping, Ma, Xiaoqi, Wei, Xiaoming, Jiao, Jianbin, Wu, Enhua, Hu, Jie
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909299208355840
author Liu, Jiajun
Wang, Yibing
Ma, Hanghang
Wu, Xiaoping
Ma, Xiaoqi
Wei, Xiaoming
Jiao, Jianbin
Wu, Enhua
Hu, Jie
author_facet Liu, Jiajun
Wang, Yibing
Ma, Hanghang
Wu, Xiaoping
Ma, Xiaoqi
Wei, Xiaoming
Jiao, Jianbin
Wu, Enhua
Hu, Jie
contents Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a challenging endeavor, especially for long videos. Due to insufficient access to large-scale high-quality video data and the excessive compression of visual features, current methods exhibit limitations in effectively processing long videos. In this paper, we introduce Kangaroo, a powerful Video LMM aimed at addressing these challenges. Confronted with issue of inadequate training data, we develop a data curation system to build a large-scale dataset with high-quality annotations for vision-language pre-training and instruction tuning. In addition, we design a curriculum training pipeline with gradually increasing resolution and number of input frames to accommodate long videos. Evaluation results demonstrate that, with 8B parameters, Kangaroo achieves state-of-the-art performance across a variety of video understanding benchmarks while exhibiting competitive results on others. Particularly, on benchmarks specialized for long videos, Kangaroo excels some larger models with over 10B parameters and proprietary models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_15542
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Liu, Jiajun
Wang, Yibing
Ma, Hanghang
Wu, Xiaoping
Ma, Xiaoqi
Wei, Xiaoming
Jiao, Jianbin
Wu, Enhua
Hu, Jie
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs). However, extending input modality of LLMs to video data remains a challenging endeavor, especially for long videos. Due to insufficient access to large-scale high-quality video data and the excessive compression of visual features, current methods exhibit limitations in effectively processing long videos. In this paper, we introduce Kangaroo, a powerful Video LMM aimed at addressing these challenges. Confronted with issue of inadequate training data, we develop a data curation system to build a large-scale dataset with high-quality annotations for vision-language pre-training and instruction tuning. In addition, we design a curriculum training pipeline with gradually increasing resolution and number of input frames to accommodate long videos. Evaluation results demonstrate that, with 8B parameters, Kangaroo achieves state-of-the-art performance across a variety of video understanding benchmarks while exhibiting competitive results on others. Particularly, on benchmarks specialized for long videos, Kangaroo excels some larger models with over 10B parameters and proprietary models.
title Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2408.15542