Learning Video Context as Interleaved Multimodal Sequences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Kevin Qinghong, Zhang, Pengchuan, Gao, Difei, Xia, Xide, Chen, Joya, Gao, Ziteng, Xie, Jinheng, Xiao, Xuhong, Shou, Mike Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916390337773568
author Lin, Kevin Qinghong
Zhang, Pengchuan
Gao, Difei
Xia, Xide
Chen, Joya
Gao, Ziteng
Xie, Jinheng
Xiao, Xuhong
Shou, Mike Zheng
author_facet Lin, Kevin Qinghong
Zhang, Pengchuan
Gao, Difei
Xia, Xide
Chen, Joya
Gao, Ziteng
Xie, Jinheng
Xiao, Xuhong
Shou, Mike Zheng
contents Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce MovieSeq, a multimodal language model developed to address the wide range of challenges in understanding video contexts. Our core idea is to represent videos as interleaved multimodal sequences (including images, plots, videos, and subtitles), either by linking external knowledge databases or using offline models (such as whisper for subtitles). Through instruction-tuning, this approach empowers the language model to interact with videos using interleaved multimodal instructions. For example, instead of solely relying on video as input, we jointly provide character photos alongside their names and dialogues, allowing the model to associate these elements and generate more comprehensive responses. To demonstrate its effectiveness, we validate MovieSeq's performance on six datasets (LVU, MAD, Movienet, CMD, TVC, MovieQA) across five settings (video classification, audio description, video-text retrieval, video captioning, and video question-answering). The code will be public at https://github.com/showlab/MovieSeq.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21757
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Video Context as Interleaved Multimodal Sequences
Lin, Kevin Qinghong
Zhang, Pengchuan
Gao, Difei
Xia, Xide
Chen, Joya
Gao, Ziteng
Xie, Jinheng
Xiao, Xuhong
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Multimedia
Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce MovieSeq, a multimodal language model developed to address the wide range of challenges in understanding video contexts. Our core idea is to represent videos as interleaved multimodal sequences (including images, plots, videos, and subtitles), either by linking external knowledge databases or using offline models (such as whisper for subtitles). Through instruction-tuning, this approach empowers the language model to interact with videos using interleaved multimodal instructions. For example, instead of solely relying on video as input, we jointly provide character photos alongside their names and dialogues, allowing the model to associate these elements and generate more comprehensive responses. To demonstrate its effectiveness, we validate MovieSeq's performance on six datasets (LVU, MAD, Movienet, CMD, TVC, MovieQA) across five settings (video classification, audio description, video-text retrieval, video captioning, and video question-answering). The code will be public at https://github.com/showlab/MovieSeq.
title Learning Video Context as Interleaved Multimodal Sequences
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2407.21757