Show Me: Unifying Instructional Image and Video Generation with Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pu, Yujiang, Huang, Zhanbo, Boddeti, Vishnu, Kong, Yu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911280349052928
author Pu, Yujiang
Huang, Zhanbo
Boddeti, Vishnu
Kong, Yu
author_facet Pu, Yujiang
Huang, Zhanbo
Boddeti, Vishnu
Kong, Yu
contents Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically treated in isolation. This separation reveals a fundamental issue: image manipulation methods overlook how actions unfold over time, while video prediction models often ignore the intended outcomes. To this end, we propose ShowMe, a unified framework that enables both tasks by selectively activating the spatial and temporal components of video diffusion models. In addition, we introduce structure and motion consistency rewards to improve structural fidelity and temporal coherence. Notably, this unification brings dual benefits: the spatial knowledge gained through video pretraining enhances contextual consistency and realism in non-rigid image edits, while the instruction-guided manipulation stage equips the model with stronger goal-oriented reasoning for video prediction. Experiments on diverse benchmarks demonstrate that our method outperforms expert models in both instructional image and video generation, highlighting the strength of video diffusion models as a unified action-object state transformer.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Pu, Yujiang
Huang, Zhanbo
Boddeti, Vishnu
Kong, Yu
Computer Vision and Pattern Recognition
Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically treated in isolation. This separation reveals a fundamental issue: image manipulation methods overlook how actions unfold over time, while video prediction models often ignore the intended outcomes. To this end, we propose ShowMe, a unified framework that enables both tasks by selectively activating the spatial and temporal components of video diffusion models. In addition, we introduce structure and motion consistency rewards to improve structural fidelity and temporal coherence. Notably, this unification brings dual benefits: the spatial knowledge gained through video pretraining enhances contextual consistency and realism in non-rigid image edits, while the instruction-guided manipulation stage equips the model with stronger goal-oriented reasoning for video prediction. Experiments on diverse benchmarks demonstrate that our method outperforms expert models in both instructional image and video generation, highlighting the strength of video diffusion models as a unified action-object state transformer.
title Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17839