TV2TV: A Unified Framework for Interleaved Language and Video Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Han, Xiaochuang, Emad, Youssef, Hall, Melissa, Nguyen, John, Padthe, Karthik, Robbins, Liam, Bar, Amir, Chen, Delong, Drozdzal, Michal, Elbayad, Maha, Hu, Yushi, Li, Shang-Wen, Roy, Sreya Dutta, Verbeek, Jakob, Wang, XuDong, Ghazvininejad, Marjan, Zettlemoyer, Luke, Dinan, Emily
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912760027152384
author Han, Xiaochuang
Emad, Youssef
Hall, Melissa
Nguyen, John
Padthe, Karthik
Robbins, Liam
Bar, Amir
Chen, Delong
Drozdzal, Michal
Elbayad, Maha
Hu, Yushi
Li, Shang-Wen
Roy, Sreya Dutta
Verbeek, Jakob
Wang, XuDong
Ghazvininejad, Marjan
Zettlemoyer, Luke
Dinan, Emily
author_facet Han, Xiaochuang
Emad, Youssef
Hall, Melissa
Nguyen, John
Padthe, Karthik
Robbins, Liam
Bar, Amir
Chen, Delong
Drozdzal, Michal
Elbayad, Maha
Hu, Yushi
Li, Shang-Wen
Roy, Sreya Dutta
Verbeek, Jakob
Wang, XuDong
Ghazvininejad, Marjan
Zettlemoyer, Luke
Dinan, Emily
contents Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new class of omni video-text models that integrate ideas from recent LM reasoning advances to address this challenge. More specifically, we present TV2TV, a unified generative modeling framework which decomposes video generation into an interleaved text and video generation process. TV2TV jointly learns language modeling (next-token prediction) and video flow matching (next-frame prediction) using a Mixture-of-Transformers (MoT) architecture. At inference time, TV2TV decides when to alternate between generating text and video frames, allowing the model to "think in words" about subsequent content before ``acting in pixels'' to produce frames. This design offloads much of the responsibility for deciding what should happen next to the language modeling tower, enabling improved visual quality and prompt alignment of generated videos. It also enables fine-grained controllability, allowing users to modify the video generation trajectory through text interventions at any point in the process. In controlled experiments on video game data, TV2TV demonstrates substantial improvements in both visual quality and controllability. TV2TV also scales to natural videos, as we show by augmenting sports videos with interleaved natural language action descriptions using vision-language models (VLMs). Training TV2TV on this corpus yields strong visual quality and prompt alignment, showcasing the model's ability to reason about and generate complex real-world action sequences. Together, these results highlight TV2TV as a promising step toward video generation with open-ended textual reasoning and control.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05103
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TV2TV: A Unified Framework for Interleaved Language and Video Generation
Han, Xiaochuang
Emad, Youssef
Hall, Melissa
Nguyen, John
Padthe, Karthik
Robbins, Liam
Bar, Amir
Chen, Delong
Drozdzal, Michal
Elbayad, Maha
Hu, Yushi
Li, Shang-Wen
Roy, Sreya Dutta
Verbeek, Jakob
Wang, XuDong
Ghazvininejad, Marjan
Zettlemoyer, Luke
Dinan, Emily
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new class of omni video-text models that integrate ideas from recent LM reasoning advances to address this challenge. More specifically, we present TV2TV, a unified generative modeling framework which decomposes video generation into an interleaved text and video generation process. TV2TV jointly learns language modeling (next-token prediction) and video flow matching (next-frame prediction) using a Mixture-of-Transformers (MoT) architecture. At inference time, TV2TV decides when to alternate between generating text and video frames, allowing the model to "think in words" about subsequent content before ``acting in pixels'' to produce frames. This design offloads much of the responsibility for deciding what should happen next to the language modeling tower, enabling improved visual quality and prompt alignment of generated videos. It also enables fine-grained controllability, allowing users to modify the video generation trajectory through text interventions at any point in the process. In controlled experiments on video game data, TV2TV demonstrates substantial improvements in both visual quality and controllability. TV2TV also scales to natural videos, as we show by augmenting sports videos with interleaved natural language action descriptions using vision-language models (VLMs). Training TV2TV on this corpus yields strong visual quality and prompt alignment, showcasing the model's ability to reason about and generate complex real-world action sequences. Together, these results highlight TV2TV as a promising step toward video generation with open-ended textual reasoning and control.
title TV2TV: A Unified Framework for Interleaved Language and Video Generation
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05103