Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Manoj, Houlsby, Neil, Hoogeboom, Emiel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917792936099840
author Kumar, Manoj
Houlsby, Neil
Hoogeboom, Emiel
author_facet Kumar, Manoj
Houlsby, Neil
Hoogeboom, Emiel
contents Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation techniques involve adapting a text-to-image model to reconstruct an input image conditioned on the same image. We first demonstrate that a diffusion model trained to reconstruct an input image from frozen embeddings, can reconstruct the image with minor variations. Second, inspired by how text-to-image models learn from web-scale text-image pairs, we explore a new pretraining strategy to generate image variations using a large collection of image pairs. Our diffusion model \textit{Semantica} receives a random (encoded) image from a webpage as conditional input and denoises another noisy random image from the same webpage. We carefully examine various design choices for the image encoder, given its crucial role in extracting relevant context from the input image. Once trained, \textit{Semantica} can adaptively generate new images from a dataset by simply using images from that dataset as input. Finally, we identify limitations in standard image consistency metrics for evaluating image variations and propose alternative metrics based on few-shot generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14857
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations
Kumar, Manoj
Houlsby, Neil
Hoogeboom, Emiel
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation techniques involve adapting a text-to-image model to reconstruct an input image conditioned on the same image. We first demonstrate that a diffusion model trained to reconstruct an input image from frozen embeddings, can reconstruct the image with minor variations. Second, inspired by how text-to-image models learn from web-scale text-image pairs, we explore a new pretraining strategy to generate image variations using a large collection of image pairs. Our diffusion model \textit{Semantica} receives a random (encoded) image from a webpage as conditional input and denoises another noisy random image from the same webpage. We carefully examine various design choices for the image encoder, given its crucial role in extracting relevant context from the input image. Once trained, \textit{Semantica} can adaptively generate new images from a dataset by simply using images from that dataset as input. Finally, we identify limitations in standard image consistency metrics for evaluating image variations and propose alternative metrics based on few-shot generation.
title Conditional Diffusion on Web-Scale Image Pairs leads to Diverse Image Variations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.14857