An Analysis of Human Alignment of Latent Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Linhardt, Lorenz, Morik, Marco, Bender, Sidney, Borras, Naima Elosegui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914713243222016
author Linhardt, Lorenz
Morik, Marco
Bender, Sidney
Borras, Naima Elosegui
author_facet Linhardt, Lorenz
Morik, Marco
Bender, Sidney
Borras, Naima Elosegui
contents Diffusion models, trained on large amounts of data, showed remarkable performance for image synthesis. They have high error consistency with humans and low texture bias when used for classification. Furthermore, prior work demonstrated the decomposability of their bottleneck layer representations into semantic directions. In this work, we analyze how well such representations are aligned to human responses on a triplet odd-one-out task. We find that despite the aforementioned observations: I) The representational alignment with humans is comparable to that of models trained only on ImageNet-1k. II) The most aligned layers of the denoiser U-Net are intermediate layers and not the bottleneck. III) Text conditioning greatly improves alignment at high noise levels, hinting at the importance of abstract textual information, especially in the early stage of generation.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08469
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Analysis of Human Alignment of Latent Diffusion Models
Linhardt, Lorenz
Morik, Marco
Bender, Sidney
Borras, Naima Elosegui
Machine Learning
Human-Computer Interaction
Diffusion models, trained on large amounts of data, showed remarkable performance for image synthesis. They have high error consistency with humans and low texture bias when used for classification. Furthermore, prior work demonstrated the decomposability of their bottleneck layer representations into semantic directions. In this work, we analyze how well such representations are aligned to human responses on a triplet odd-one-out task. We find that despite the aforementioned observations: I) The representational alignment with humans is comparable to that of models trained only on ImageNet-1k. II) The most aligned layers of the denoiser U-Net are intermediate layers and not the bottleneck. III) Text conditioning greatly improves alignment at high noise levels, hinting at the importance of abstract textual information, especially in the early stage of generation.
title An Analysis of Human Alignment of Latent Diffusion Models
topic Machine Learning
Human-Computer Interaction
url https://arxiv.org/abs/2403.08469