Everything is a Video: Unifying Modalities through Next-Frame Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hudson, G. Thomas, Slack, Dean, Winterbottom, Thomas, Sterling, Jamie, Xiao, Chenghao, Shentu, Junjie, Moubayed, Noura Al
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911078823231488
author Hudson, G. Thomas
Slack, Dean
Winterbottom, Thomas
Sterling, Jamie
Xiao, Chenghao
Shentu, Junjie
Moubayed, Noura Al
author_facet Hudson, G. Thomas
Slack, Dean
Winterbottom, Thomas
Sterling, Jamie
Xiao, Chenghao
Shentu, Junjie
Moubayed, Noura Al
contents Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation. Traditional approaches rely on modality-specific encoders and late fusion techniques, which can hinder scalability and flexibility when adapting to new tasks or modalities. To address these limitations, we introduce a novel framework that extends the concept of task reformulation beyond natural language processing (NLP) to multimodal learning. We propose to reformulate diverse multimodal tasks into a unified next-frame prediction problem, allowing a single model to handle different modalities without modality-specific components. This method treats all inputs and outputs as sequential frames in a video, enabling seamless integration of modalities and effective knowledge transfer across tasks. Our approach is evaluated on a range of tasks, including text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text, demonstrating the model's ability to generalize across modalities with minimal adaptation. We show that task reformulation can significantly simplify multimodal model design across various tasks, laying the groundwork for more generalized multimodal foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10503
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Everything is a Video: Unifying Modalities through Next-Frame Prediction
Hudson, G. Thomas
Slack, Dean
Winterbottom, Thomas
Sterling, Jamie
Xiao, Chenghao
Shentu, Junjie
Moubayed, Noura Al
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Multimodal learning, which involves integrating information from various modalities such as text, images, audio, and video, is pivotal for numerous complex tasks like visual question answering, cross-modal retrieval, and caption generation. Traditional approaches rely on modality-specific encoders and late fusion techniques, which can hinder scalability and flexibility when adapting to new tasks or modalities. To address these limitations, we introduce a novel framework that extends the concept of task reformulation beyond natural language processing (NLP) to multimodal learning. We propose to reformulate diverse multimodal tasks into a unified next-frame prediction problem, allowing a single model to handle different modalities without modality-specific components. This method treats all inputs and outputs as sequential frames in a video, enabling seamless integration of modalities and effective knowledge transfer across tasks. Our approach is evaluated on a range of tasks, including text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text, demonstrating the model's ability to generalize across modalities with minimal adaptation. We show that task reformulation can significantly simplify multimodal model design across various tasks, laying the groundwork for more generalized multimodal foundation models.
title Everything is a Video: Unifying Modalities through Next-Frame Prediction
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.10503