Learning Video Context as Interleaved Multimodal Sequences
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916390337773568 |
|---|---|
| author | Lin, Kevin Qinghong Zhang, Pengchuan Gao, Difei Xia, Xide Chen, Joya Gao, Ziteng Xie, Jinheng Xiao, Xuhong Shou, Mike Zheng |
| author_facet | Lin, Kevin Qinghong Zhang, Pengchuan Gao, Difei Xia, Xide Chen, Joya Gao, Ziteng Xie, Jinheng Xiao, Xuhong Shou, Mike Zheng |
| contents | Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce MovieSeq, a multimodal language model developed to address the wide range of challenges in understanding video contexts. Our core idea is to represent videos as interleaved multimodal sequences (including images, plots, videos, and subtitles), either by linking external knowledge databases or using offline models (such as whisper for subtitles). Through instruction-tuning, this approach empowers the language model to interact with videos using interleaved multimodal instructions. For example, instead of solely relying on video as input, we jointly provide character photos alongside their names and dialogues, allowing the model to associate these elements and generate more comprehensive responses. To demonstrate its effectiveness, we validate MovieSeq's performance on six datasets (LVU, MAD, Movienet, CMD, TVC, MovieQA) across five settings (video classification, audio description, video-text retrieval, video captioning, and video question-answering). The code will be public at https://github.com/showlab/MovieSeq. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_21757 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning Video Context as Interleaved Multimodal Sequences Lin, Kevin Qinghong Zhang, Pengchuan Gao, Difei Xia, Xide Chen, Joya Gao, Ziteng Xie, Jinheng Xiao, Xuhong Shou, Mike Zheng Computer Vision and Pattern Recognition Multimedia Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce MovieSeq, a multimodal language model developed to address the wide range of challenges in understanding video contexts. Our core idea is to represent videos as interleaved multimodal sequences (including images, plots, videos, and subtitles), either by linking external knowledge databases or using offline models (such as whisper for subtitles). Through instruction-tuning, this approach empowers the language model to interact with videos using interleaved multimodal instructions. For example, instead of solely relying on video as input, we jointly provide character photos alongside their names and dialogues, allowing the model to associate these elements and generate more comprehensive responses. To demonstrate its effectiveness, we validate MovieSeq's performance on six datasets (LVU, MAD, Movienet, CMD, TVC, MovieQA) across five settings (video classification, audio description, video-text retrieval, video captioning, and video question-answering). The code will be public at https://github.com/showlab/MovieSeq. |
| title | Learning Video Context as Interleaved Multimodal Sequences |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2407.21757 |