SliderSpace: Decomposing the Visual Capabilities of Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gandikota, Rohit, Wu, Zongze, Zhang, Richard, Bau, David, Shechtman, Eli, Kolkin, Nick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917910599958528
author Gandikota, Rohit
Wu, Zongze
Zhang, Richard
Bau, David
Shechtman, Eli
Kolkin, Nick
author_facet Gandikota, Rohit
Wu, Zongze
Zhang, Richard
Bau, David
Shechtman, Eli
Kolkin, Nick
contents We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to specify attributes for each edit direction individually, SliderSpace discovers multiple interpretable and diverse directions simultaneously from a single text prompt. Each direction is trained as a low-rank adaptor, enabling compositional control and the discovery of surprising possibilities in the model's latent space. Through extensive experiments on state-of-the-art diffusion models, we demonstrate SliderSpace's effectiveness across three applications: concept decomposition, artistic style exploration, and diversity enhancement. Our quantitative evaluation shows that SliderSpace-discovered directions decompose the visual structure of model's knowledge effectively, offering insights into the latent capabilities encoded within diffusion models. User studies further validate that our method produces more diverse and useful variations compared to baselines. Our code, data and trained weights are available at https://sliderspace.baulab.info
format Preprint
id arxiv_https___arxiv_org_abs_2502_01639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
Gandikota, Rohit
Wu, Zongze
Zhang, Richard
Bau, David
Shechtman, Eli
Kolkin, Nick
Computer Vision and Pattern Recognition
Graphics
Machine Learning
We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to specify attributes for each edit direction individually, SliderSpace discovers multiple interpretable and diverse directions simultaneously from a single text prompt. Each direction is trained as a low-rank adaptor, enabling compositional control and the discovery of surprising possibilities in the model's latent space. Through extensive experiments on state-of-the-art diffusion models, we demonstrate SliderSpace's effectiveness across three applications: concept decomposition, artistic style exploration, and diversity enhancement. Our quantitative evaluation shows that SliderSpace-discovered directions decompose the visual structure of model's knowledge effectively, offering insights into the latent capabilities encoded within diffusion models. User studies further validate that our method produces more diverse and useful variations compared to baselines. Our code, data and trained weights are available at https://sliderspace.baulab.info
title SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
topic Computer Vision and Pattern Recognition
Graphics
Machine Learning
url https://arxiv.org/abs/2502.01639