DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Byun, Dongnam, Park, Jungwon, Ko, Jungmin, Choi, Changin, Rhee, Wonjong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910080698417152
author Byun, Dongnam
Park, Jungwon
Ko, Jungmin
Choi, Changin
Rhee, Wonjong
author_facet Byun, Dongnam
Park, Jungwon
Ko, Jungmin
Choi, Changin
Rhee, Wonjong
contents Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14376
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
Byun, Dongnam
Park, Jungwon
Ko, Jungmin
Choi, Changin
Rhee, Wonjong
Computer Vision and Pattern Recognition
Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.
title DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14376