GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Souček, Tomáš, Damen, Dima, Wray, Michael, Laptev, Ivan, Sivic, Josef
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916188282421248
author Souček, Tomáš
Damen, Dima
Wray, Michael
Laptev, Ivan
Sivic, Josef
author_facet Souček, Tomáš
Damen, Dima
Wray, Michael
Laptev, Ivan
Sivic, Josef
contents We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2312_07322
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
Souček, Tomáš
Damen, Dima
Wray, Michael
Laptev, Ivan
Sivic, Josef
Computer Vision and Pattern Recognition
We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.
title GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.07322