Pathways on the Image Manifold: Image Editing via Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rotstein, Noam, Yona, Gal, Silver, Daniel, Velich, Roy, Bensaïd, David, Kimmel, Ron
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910885042192384
author Rotstein, Noam
Yona, Gal
Silver, Daniel
Velich, Roy
Bensaïd, David
Kimmel, Ron
author_facet Rotstein, Noam
Yona, Gal
Silver, Daniel
Velich, Roy
Bensaïd, David
Kimmel, Ron
contents Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently compromise fidelity by altering key elements of the original image. Simultaneously, video generation has made remarkable strides, with models that effectively function as consistent and continuous world simulators. In this paper, we propose merging these two fields by utilizing image-to-video models for image editing. We reformulate image editing as a temporal process, using pretrained video models to create smooth transitions from the original image to the desired edit. This approach traverses the image manifold continuously, ensuring consistent edits while preserving the original image's key aspects. Our approach achieves state-of-the-art results on text-based image editing, demonstrating significant improvements in both edit accuracy and image preservation. Visit our project page at https://rotsteinnoam.github.io/Frame2Frame.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pathways on the Image Manifold: Image Editing via Video Generation
Rotstein, Noam
Yona, Gal
Silver, Daniel
Velich, Roy
Bensaïd, David
Kimmel, Ron
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently compromise fidelity by altering key elements of the original image. Simultaneously, video generation has made remarkable strides, with models that effectively function as consistent and continuous world simulators. In this paper, we propose merging these two fields by utilizing image-to-video models for image editing. We reformulate image editing as a temporal process, using pretrained video models to create smooth transitions from the original image to the desired edit. This approach traverses the image manifold continuously, ensuring consistent edits while preserving the original image's key aspects. Our approach achieves state-of-the-art results on text-based image editing, demonstrating significant improvements in both edit accuracy and image preservation. Visit our project page at https://rotsteinnoam.github.io/Frame2Frame.
title Pathways on the Image Manifold: Image Editing via Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.16819