Steer Away From Mode Collisions: Improving Composition In Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dutta, Debottam, Chen, Jianchong, Rajagopalan, Rajalaxmi, Wei, Yu-Lin, Choudhury, Romit Roy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915860543700992
author Dutta, Debottam
Chen, Jianchong
Rajagopalan, Rajalaxmi
Wei, Yu-Lin
Choudhury, Romit Roy
author_facet Dutta, Debottam
Chen, Jianchong
Rajagopalan, Rajalaxmi
Wei, Yu-Lin
Choudhury, Romit Roy
contents We propose to improve multi-concept prompt fidelity in text-to-image diffusion models. We begin with common failure cases - prompts like "a cat and a dog" that sometimes yields images where one concept is missing, faint, or colliding awkwardly with another. We hypothesize that this happens when the diffusion model drifts into mixed modes that over-emphasize a single concept it learned strongly during training. Instead of re-training, we introduce a corrective sampling strategy that steers away from regions where the joint prompt behavior overlaps too strongly with any single concept in the prompt. The goal is to steer towards "pure" joint modes where all concepts can coexist with balanced visual presence. We further show that existing multi-concept guidance schemes can operate in unstable weight regimes that amplify imbalance; we characterize favorable regions and adapt sampling to remain within them. Our approach, CO3, is plug-and-play, requires no model tuning, and complements standard classifier-free guidance. Experiments on diverse multi-concept prompts indicate improvements in concept coverage, balance and robustness, with fewer dropped or distorted concepts compared to standard baselines and prior compositional methods. Results suggest that lightweight corrective guidance can substantially mitigate brittle semantic alignment behavior in modern diffusion systems. Code is available at https://github.com/debottam-dutta7/co3
format Preprint
id arxiv_https___arxiv_org_abs_2509_25940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steer Away From Mode Collisions: Improving Composition In Diffusion Models
Dutta, Debottam
Chen, Jianchong
Rajagopalan, Rajalaxmi
Wei, Yu-Lin
Choudhury, Romit Roy
Computer Vision and Pattern Recognition
Machine Learning
We propose to improve multi-concept prompt fidelity in text-to-image diffusion models. We begin with common failure cases - prompts like "a cat and a dog" that sometimes yields images where one concept is missing, faint, or colliding awkwardly with another. We hypothesize that this happens when the diffusion model drifts into mixed modes that over-emphasize a single concept it learned strongly during training. Instead of re-training, we introduce a corrective sampling strategy that steers away from regions where the joint prompt behavior overlaps too strongly with any single concept in the prompt. The goal is to steer towards "pure" joint modes where all concepts can coexist with balanced visual presence. We further show that existing multi-concept guidance schemes can operate in unstable weight regimes that amplify imbalance; we characterize favorable regions and adapt sampling to remain within them. Our approach, CO3, is plug-and-play, requires no model tuning, and complements standard classifier-free guidance. Experiments on diverse multi-concept prompts indicate improvements in concept coverage, balance and robustness, with fewer dropped or distorted concepts compared to standard baselines and prior compositional methods. Results suggest that lightweight corrective guidance can substantially mitigate brittle semantic alignment behavior in modern diffusion systems. Code is available at https://github.com/debottam-dutta7/co3
title Steer Away From Mode Collisions: Improving Composition In Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.25940