An Analysis of Human Alignment of Latent Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914713243222016 |
|---|---|
| author | Linhardt, Lorenz Morik, Marco Bender, Sidney Borras, Naima Elosegui |
| author_facet | Linhardt, Lorenz Morik, Marco Bender, Sidney Borras, Naima Elosegui |
| contents | Diffusion models, trained on large amounts of data, showed remarkable performance for image synthesis. They have high error consistency with humans and low texture bias when used for classification. Furthermore, prior work demonstrated the decomposability of their bottleneck layer representations into semantic directions. In this work, we analyze how well such representations are aligned to human responses on a triplet odd-one-out task. We find that despite the aforementioned observations: I) The representational alignment with humans is comparable to that of models trained only on ImageNet-1k. II) The most aligned layers of the denoiser U-Net are intermediate layers and not the bottleneck. III) Text conditioning greatly improves alignment at high noise levels, hinting at the importance of abstract textual information, especially in the early stage of generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_08469 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | An Analysis of Human Alignment of Latent Diffusion Models Linhardt, Lorenz Morik, Marco Bender, Sidney Borras, Naima Elosegui Machine Learning Human-Computer Interaction Diffusion models, trained on large amounts of data, showed remarkable performance for image synthesis. They have high error consistency with humans and low texture bias when used for classification. Furthermore, prior work demonstrated the decomposability of their bottleneck layer representations into semantic directions. In this work, we analyze how well such representations are aligned to human responses on a triplet odd-one-out task. We find that despite the aforementioned observations: I) The representational alignment with humans is comparable to that of models trained only on ImageNet-1k. II) The most aligned layers of the denoiser U-Net are intermediate layers and not the bottleneck. III) Text conditioning greatly improves alignment at high noise levels, hinting at the importance of abstract textual information, especially in the early stage of generation. |
| title | An Analysis of Human Alignment of Latent Diffusion Models |
| topic | Machine Learning Human-Computer Interaction |
| url | https://arxiv.org/abs/2403.08469 |