PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Yanjie, He, Qingdong, Jiang, Zhengkai, Xu, Pengcheng, Wang, Chaoyi, Peng, Jinlong, Wang, Haoxuan, Cao, Yun, Gan, Zhenye, Chi, Mingmin, Peng, Bo, Wang, Yabiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913153502150656
author Pan, Yanjie
He, Qingdong
Jiang, Zhengkai
Xu, Pengcheng
Wang, Chaoyi
Peng, Jinlong
Wang, Haoxuan
Cao, Yun
Gan, Zhenye
Chi, Mingmin
Peng, Bo
Wang, Yabiao
author_facet Pan, Yanjie
He, Qingdong
Jiang, Zhengkai
Xu, Pengcheng
Wang, Chaoyi
Peng, Jinlong
Wang, Haoxuan
Cao, Yun
Gan, Zhenye
Chi, Mingmin
Peng, Bo
Wang, Yabiao
contents Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously preserving semantic fidelity across multiple heterogeneous control signals while maintaining high visual quality, where they employ separate control branches that often introduce conflicting guidance during the denoising process, leading to structural distortions and artifacts in generated images. To address this issue, we present PixelPonder, a novel unified control framework, which allows for effective control of multiple visual conditions under a single control structure. Specifically, we design a patch-level adaptive condition selection mechanism that dynamically prioritizes spatially relevant control signals at the sub-region level, enabling precise local guidance without global interference. Additionally, a time-aware control injection scheme is deployed to modulate condition influence according to denoising timesteps, progressively transitioning from structural preservation to texture refinement and fully utilizing the control information from different categories to promote more harmonious image generation. Extensive experiments demonstrate that PixelPonder surpasses previous methods across different benchmark datasets, showing superior improvement in spatial alignment accuracy while maintaining high textual semantic consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
Pan, Yanjie
He, Qingdong
Jiang, Zhengkai
Xu, Pengcheng
Wang, Chaoyi
Peng, Jinlong
Wang, Haoxuan
Cao, Yun
Gan, Zhenye
Chi, Mingmin
Peng, Bo
Wang, Yabiao
Computer Vision and Pattern Recognition
Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously preserving semantic fidelity across multiple heterogeneous control signals while maintaining high visual quality, where they employ separate control branches that often introduce conflicting guidance during the denoising process, leading to structural distortions and artifacts in generated images. To address this issue, we present PixelPonder, a novel unified control framework, which allows for effective control of multiple visual conditions under a single control structure. Specifically, we design a patch-level adaptive condition selection mechanism that dynamically prioritizes spatially relevant control signals at the sub-region level, enabling precise local guidance without global interference. Additionally, a time-aware control injection scheme is deployed to modulate condition influence according to denoising timesteps, progressively transitioning from structural preservation to texture refinement and fully utilizing the control information from different categories to promote more harmonious image generation. Extensive experiments demonstrate that PixelPonder surpasses previous methods across different benchmark datasets, showing superior improvement in spatial alignment accuracy while maintaining high textual semantic consistency.
title PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06684