PARASOL: Parametric Style Control for Diffusion Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911861784444928 |
|---|---|
| author | Tarrés, Gemma Canet Ruta, Dan Bui, Tu Collomosse, John |
| author_facet | Tarrés, Gemma Canet Ruta, Dan Bui, Tu Collomosse, John |
| contents | We propose PARASOL, a multi-modal synthesis model that enables disentangled, parametric control of the visual style of the image by jointly conditioning synthesis on both content and a fine-grained visual style embedding. We train a latent diffusion model (LDM) using specific losses for each modality and adapt the classifier-free guidance for encouraging disentangled control over independent content and style modalities at inference time. We leverage auxiliary semantic and style-based search to create training triplets for supervision of the LDM, ensuring complementarity of content and style cues. PARASOL shows promise for enabling nuanced control over visual style in diffusion models for image creation and stylization, as well as generative search where text-based search results may be adapted to more closely match user intent by interpolating both content and style descriptors. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2303_06464 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | PARASOL: Parametric Style Control for Diffusion Image Synthesis Tarrés, Gemma Canet Ruta, Dan Bui, Tu Collomosse, John Computer Vision and Pattern Recognition We propose PARASOL, a multi-modal synthesis model that enables disentangled, parametric control of the visual style of the image by jointly conditioning synthesis on both content and a fine-grained visual style embedding. We train a latent diffusion model (LDM) using specific losses for each modality and adapt the classifier-free guidance for encouraging disentangled control over independent content and style modalities at inference time. We leverage auxiliary semantic and style-based search to create training triplets for supervision of the LDM, ensuring complementarity of content and style cues. PARASOL shows promise for enabling nuanced control over visual style in diffusion models for image creation and stylization, as well as generative search where text-based search results may be adapted to more closely match user intent by interpolating both content and style descriptors. |
| title | PARASOL: Parametric Style Control for Diffusion Image Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2303.06464 |