Synthetic Curriculum Reinforces Compositional Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Shijian, Fu, Runhao, Zhao, Siyi, Zhan, Qingqin, Wang, Xingjian, Jin, Jiarui, Lu, Yuan, Wu, Hanqian, Chen, Cunjian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911283041796096
author Wang, Shijian
Fu, Runhao
Zhao, Siyi
Zhan, Qingqin
Wang, Xingjian
Jin, Jiarui
Lu, Yuan
Wu, Hanqian
Chen, Cunjian
author_facet Wang, Shijian
Fu, Runhao
Zhao, Siyi
Zhan, Qingqin
Wang, Xingjian
Jin, Jiarui
Lu, Yuan
Wu, Hanqian
Chen, Cunjian
contents Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhibit diverse attributes as well as intricate spatial and semantic relationships, demanding both precise object placement and coherent inter-object interactions. In this paper, we propose a novel compositional curriculum reinforcement learning framework named CompGen that addresses compositional weakness in existing T2I models. Specifically, we leverage scene graphs to establish a novel difficulty criterion for compositional ability and develop a corresponding adaptive Markov Chain Monte Carlo graph sampling algorithm. This difficulty-aware approach enables the synthesis of training curriculum data that progressively optimize T2I models through reinforcement learning. We integrate our curriculum learning approach into Group Relative Policy Optimization (GRPO) and investigate different curriculum scheduling strategies. Our experiments reveal that CompGen exhibits distinct scaling curves under different curriculum scheduling strategies, with easy-to-hard and Gaussian sampling strategies yielding superior scaling performance compared to random sampling. Extensive experiments demonstrate that CompGen significantly enhances compositional generation capabilities for both diffusion-based and auto-regressive T2I models, highlighting its effectiveness in improving the compositional T2I generation systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18378
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
Wang, Shijian
Fu, Runhao
Zhao, Siyi
Zhan, Qingqin
Wang, Xingjian
Jin, Jiarui
Lu, Yuan
Wu, Hanqian
Chen, Cunjian
Computer Vision and Pattern Recognition
Text-to-Image (T2I) generation has long been an open problem, with compositional synthesis remaining particularly challenging. This task requires accurate rendering of complex scenes containing multiple objects that exhibit diverse attributes as well as intricate spatial and semantic relationships, demanding both precise object placement and coherent inter-object interactions. In this paper, we propose a novel compositional curriculum reinforcement learning framework named CompGen that addresses compositional weakness in existing T2I models. Specifically, we leverage scene graphs to establish a novel difficulty criterion for compositional ability and develop a corresponding adaptive Markov Chain Monte Carlo graph sampling algorithm. This difficulty-aware approach enables the synthesis of training curriculum data that progressively optimize T2I models through reinforcement learning. We integrate our curriculum learning approach into Group Relative Policy Optimization (GRPO) and investigate different curriculum scheduling strategies. Our experiments reveal that CompGen exhibits distinct scaling curves under different curriculum scheduling strategies, with easy-to-hard and Gaussian sampling strategies yielding superior scaling performance compared to random sampling. Extensive experiments demonstrate that CompGen significantly enhances compositional generation capabilities for both diffusion-based and auto-regressive T2I models, highlighting its effectiveness in improving the compositional T2I generation systems.
title Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18378