CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Gaoyang, Fu, Bingtao, Fan, Qingnan, Zhang, Qi, Liu, Runxing, Gu, Hong, Zhang, Huaqi, Liu, Xinguo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916915393331200
author Zhang, Gaoyang
Fu, Bingtao
Fan, Qingnan
Zhang, Qi
Liu, Runxing
Gu, Hong
Zhang, Huaqi
Liu, Xinguo
author_facet Zhang, Gaoyang
Fu, Bingtao
Fan, Qingnan
Zhang, Qi
Liu, Runxing
Gu, Hong
Zhang, Huaqi
Liu, Xinguo
contents Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
Zhang, Gaoyang
Fu, Bingtao
Fan, Qingnan
Zhang, Qi
Liu, Runxing
Gu, Hong
Zhang, Huaqi
Liu, Xinguo
Computer Vision and Pattern Recognition
Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS.
title CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13195