How to Blend Concepts in Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Olearo, Lorenzo, Longari, Giorgio, Melzi, Simone, Raganato, Alessandro, Peñaloza, Rafael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909321816702976
author Olearo, Lorenzo
Longari, Giorgio
Melzi, Simone
Raganato, Alessandro
Peñaloza, Rafael
author_facet Olearo, Lorenzo
Longari, Giorgio
Melzi, Simone
Raganato, Alessandro
Peñaloza, Rafael
contents For the last decade, there has been a push to use multi-dimensional (latent) spaces to represent concepts; and yet how to manipulate these concepts or reason with them remains largely unclear. Some recent methods exploit multiple latent representations and their connection, making this research question even more entangled. Our goal is to understand how operations in the latent space affect the underlying concepts. To that end, we explore the task of concept blending through diffusion models. Diffusion models are based on a connection between a latent representation of textual prompts and a latent space that enables image reconstruction and generation. This task allows us to try different text-based combination strategies, and evaluate easily through a visual analysis. Our conclusion is that concept blending through space manipulation is possible, although the best strategy depends on the context of the blend.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14280
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How to Blend Concepts in Diffusion Models
Olearo, Lorenzo
Longari, Giorgio
Melzi, Simone
Raganato, Alessandro
Peñaloza, Rafael
Computer Vision and Pattern Recognition
Artificial Intelligence
For the last decade, there has been a push to use multi-dimensional (latent) spaces to represent concepts; and yet how to manipulate these concepts or reason with them remains largely unclear. Some recent methods exploit multiple latent representations and their connection, making this research question even more entangled. Our goal is to understand how operations in the latent space affect the underlying concepts. To that end, we explore the task of concept blending through diffusion models. Diffusion models are based on a connection between a latent representation of textual prompts and a latent space that enables image reconstruction and generation. This task allows us to try different text-based combination strategies, and evaluate easily through a visual analysis. Our conclusion is that concept blending through space manipulation is possible, although the best strategy depends on the context of the blend.
title How to Blend Concepts in Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.14280