Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhenyu, Xie, Enze, Li, Aoxue, Wang, Zhongdao, Liu, Xihui, Li, Zhenguo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023)
by: Huang, Kaiyi, et al.
Published: (2023)
Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion
by: Li, Aoxue, et al.
Published: (2024)
by: Li, Aoxue, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model
by: Yi, Mingyang, et al.
Published: (2024)
by: Yi, Mingyang, et al.
Published: (2024)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
by: Sun, Kaiyue, et al.
Published: (2024)
by: Sun, Kaiyue, et al.
Published: (2024)
MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
by: Wang, Yanhui, et al.
Published: (2023)
by: Wang, Yanhui, et al.
Published: (2023)
Efficient Transferability Assessment for Selection of Pre-trained Detectors
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding
by: Teng, Yao, et al.
Published: (2024)
by: Teng, Yao, et al.
Published: (2024)
DivCon: Divide and Conquer for Complex Numerical and Spatial Reasoning in Text-to-Image Generation
by: Jia, Yuhao, et al.
Published: (2024)
by: Jia, Yuhao, et al.
Published: (2024)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
by: Wang, Wenbin, et al.
Published: (2024)
by: Wang, Wenbin, et al.
Published: (2024)
Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
SJD++: Improved Speculative Jacobi Decoding for Training-free Acceleration of Discrete Auto-regressive Text-to-Image Generation
by: Teng, Yao, et al.
Published: (2025)
by: Teng, Yao, et al.
Published: (2025)
Personalized Text-to-Image Generation with Auto-Regressive Models
by: Sun, Kaiyue, et al.
Published: (2025)
by: Sun, Kaiyue, et al.
Published: (2025)
Memory Consistency Guided Divide-and-Conquer Learning for Generalized Category Discovery
by: Tu, Yuanpeng, et al.
Published: (2024)
by: Tu, Yuanpeng, et al.
Published: (2024)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
by: Teng, Yao, et al.
Published: (2025)
by: Teng, Yao, et al.
Published: (2025)
Divide-and-Conquer Approach to Holistic Cognition in High-Similarity Contexts with Limited Data
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
by: Wang, Yunnan, et al.
Published: (2024)
by: Wang, Yunnan, et al.
Published: (2024)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
by: Gan, Chaofan, et al.
Published: (2024)
by: Gan, Chaofan, et al.
Published: (2024)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
Semantics versus Identity: A Divide-and-Conquer Approach towards Adjustable Medical Image De-Identification
by: Tian, Yuan, et al.
Published: (2025)
by: Tian, Yuan, et al.
Published: (2025)
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
by: Wang, Tianqi, et al.
Published: (2024)
by: Wang, Tianqi, et al.
Published: (2024)
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
by: Ye, Zhipeng, et al.
Published: (2026)
by: Ye, Zhipeng, et al.
Published: (2026)
Divide and Conquer: Grounding a Bleeding Areas in Gastrointestinal Image with Two-Stage Model
by: Lin, Yu-Fan, et al.
Published: (2024)
by: Lin, Yu-Fan, et al.
Published: (2024)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
by: Huang, Kaiyi, et al.
Published: (2024)
by: Huang, Kaiyi, et al.
Published: (2024)
DCA: Dividing and Conquering Amnesia in Incremental Object Detection
by: Zhang, Aoting, et al.
Published: (2025)
by: Zhang, Aoting, et al.
Published: (2025)
ShadowHack: Hacking Shadows via Luminance-Color Divide and Conquer
by: Hu, Jin, et al.
Published: (2024)
by: Hu, Jin, et al.
Published: (2024)
Divide and Conquer Self-Supervised Learning for High-Content Imaging
by: Farndale, Lucas, et al.
Published: (2025)
by: Farndale, Lucas, et al.
Published: (2025)
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
by: Xiao, Junyuan, et al.
Published: (2026)
by: Xiao, Junyuan, et al.
Published: (2026)
Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection
by: Kang, Xiaolu, et al.
Published: (2026)
by: Kang, Xiaolu, et al.
Published: (2026)
Animate124: Animating One Image to 4D Dynamic Scene
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Similar Items
-
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
by: Wang, Zhenyu, et al.
Published: (2024) -
T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation
by: Huang, Kaiyi, et al.
Published: (2023) -
Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion
by: Li, Aoxue, et al.
Published: (2024) -
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024) -
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)