Saved in:
Bibliographic Details
Main Authors: Ferrara, Alfio, Picascia, Sergio, Rocchetti, Elisabetta
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.23313
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910240134397952
author Ferrara, Alfio
Picascia, Sergio
Rocchetti, Elisabetta
author_facet Ferrara, Alfio
Picascia, Sergio
Rocchetti, Elisabetta
contents Text-to-image diffusion models have demonstrated remarkable capabilities in generating artistic content by learning from billions of images, including popular artworks. However, the fundamental question of how these models internally represent concepts, such as content and style in paintings, remains unexplored. Traditional computer vision assumes content and style are orthogonal, but diffusion models receive no explicit guidance about this distinction during training. In this work, we investigate how transformer-based text-to-image diffusion models encode content and style concepts when generating artworks. We leverage cross-attention heatmaps to attribute pixels in generated images to specific prompt tokens, enabling us to isolate image regions influenced by content-describing versus style-describing tokens. Our findings reveal that diffusion models demonstrate varying degrees of content-style separation depending on the specific artistic prompt and style requested. In many cases, content tokens primarily influence object-related regions while style tokens affect background and texture areas, suggesting an emergent understanding of the content-style distinction. These insights contribute to our understanding of how large-scale generative models internally represent complex artistic concepts without explicit supervision. We share the code and dataset, together with an exploratory tool for visualizing attention maps at https://github.com/umilISLab/artistic-prompt-interpretation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23313
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models
Ferrara, Alfio
Picascia, Sergio
Rocchetti, Elisabetta
Computer Vision and Pattern Recognition
Text-to-image diffusion models have demonstrated remarkable capabilities in generating artistic content by learning from billions of images, including popular artworks. However, the fundamental question of how these models internally represent concepts, such as content and style in paintings, remains unexplored. Traditional computer vision assumes content and style are orthogonal, but diffusion models receive no explicit guidance about this distinction during training. In this work, we investigate how transformer-based text-to-image diffusion models encode content and style concepts when generating artworks. We leverage cross-attention heatmaps to attribute pixels in generated images to specific prompt tokens, enabling us to isolate image regions influenced by content-describing versus style-describing tokens. Our findings reveal that diffusion models demonstrate varying degrees of content-style separation depending on the specific artistic prompt and style requested. In many cases, content tokens primarily influence object-related regions while style tokens affect background and texture areas, suggesting an emergent understanding of the content-style distinction. These insights contribute to our understanding of how large-scale generative models internally represent complex artistic concepts without explicit supervision. We share the code and dataset, together with an exploratory tool for visualizing attention maps at https://github.com/umilISLab/artistic-prompt-interpretation.
title The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.23313