VideoOrion: Tokenizing Object Dynamics in Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Yicheng, Li, Yijiang, Zhang, Wanpeng, Luo, Hao, Yue, Zihao, Zheng, Sipeng, Lu, Zongqing
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912279038001152
author Feng, Yicheng
Li, Yijiang
Zhang, Wanpeng
Luo, Hao
Yue, Zihao
Zheng, Sipeng
Lu, Zongqing
author_facet Feng, Yicheng
Li, Yijiang
Zhang, Wanpeng
Luo, Hao
Yue, Zihao
Zheng, Sipeng
Lu, Zongqing
contents We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating spatial-temporal object features. Our method addresses the persistent challenge in Video-LLMs of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs. Compared to prior methods which resort to downsampling the original video or aggregating visual tokens using resamplers, leading to information loss and entangled semantics, VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations but also enables explicit object modeling of video content with minimal computational cost. Moreover, the introduced object tokens naturally allow VideoOrion to accomplish video-based referring tasks. Experimental results show that VideoOrion can learn to make good use of the object tokens, and achieves competitive results on both general video question answering and video-based referring benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16156
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VideoOrion: Tokenizing Object Dynamics in Videos
Feng, Yicheng
Li, Yijiang
Zhang, Wanpeng
Luo, Hao
Yue, Zihao
Zheng, Sipeng
Lu, Zongqing
Computer Vision and Pattern Recognition
Machine Learning
We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating spatial-temporal object features. Our method addresses the persistent challenge in Video-LLMs of efficiently compressing high-dimensional video data into semantic tokens that are comprehensible to LLMs. Compared to prior methods which resort to downsampling the original video or aggregating visual tokens using resamplers, leading to information loss and entangled semantics, VideoOrion not only offers a more natural and efficient way to derive compact, disentangled semantic representations but also enables explicit object modeling of video content with minimal computational cost. Moreover, the introduced object tokens naturally allow VideoOrion to accomplish video-based referring tasks. Experimental results show that VideoOrion can learn to make good use of the object tokens, and achieves competitive results on both general video question answering and video-based referring benchmarks.
title VideoOrion: Tokenizing Object Dynamics in Videos
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.16156