Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shijian, Fu, Runhao, Zhao, Siyi, Zhan, Qingqin, Wang, Xingjian, Jin, Jiarui, Lu, Yuan, Wu, Hanqian, Chen, Cunjian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024)
by: Wang, Shijian, et al.
Published: (2024)
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
by: Wang, Shijian, et al.
Published: (2026)
by: Wang, Shijian, et al.
Published: (2026)
Training-free Stylized Text-to-Image Generation with Fast Inference
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
by: Yan, Zexuan, et al.
Published: (2026)
by: Yan, Zexuan, et al.
Published: (2026)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
by: Sun, Haoyuan, et al.
Published: (2026)
by: Sun, Haoyuan, et al.
Published: (2026)
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
by: Gao, Qiang, et al.
Published: (2026)
by: Gao, Qiang, et al.
Published: (2026)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
by: Wang, Juntong, et al.
Published: (2025)
by: Wang, Juntong, et al.
Published: (2025)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
CSGO: Content-Style Composition in Text-to-Image Generation
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
FT2TF: First-Person Statement Text-To-Talking Face Generation
by: Diao, Xingjian, et al.
Published: (2023)
by: Diao, Xingjian, et al.
Published: (2023)
Progressive Compositionality in Text-to-Image Generative Models
by: Han, Evans Xu, et al.
Published: (2024)
by: Han, Evans Xu, et al.
Published: (2024)
LEO: Generative Latent Image Animator for Human Video Synthesis
by: Wang, Yaohui, et al.
Published: (2023)
by: Wang, Yaohui, et al.
Published: (2023)
Weakly Supervised Monocular 3D Detection with a Single-View Image
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis
by: Su, Tongtong, et al.
Published: (2025)
by: Su, Tongtong, et al.
Published: (2025)
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Versatile Transition Generation with Image-to-Video Diffusion
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Rectifying Latent Space for Generative Single-Image Reflection Removal
by: Li, Mingjia, et al.
Published: (2025)
by: Li, Mingjia, et al.
Published: (2025)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
Generating Intermediate Representations for Compositional Text-To-Image Generation
by: Galun, Ran, et al.
Published: (2024)
by: Galun, Ran, et al.
Published: (2024)
POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation
by: Fan, Yaohou, et al.
Published: (2026)
by: Fan, Yaohou, et al.
Published: (2026)
AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
by: He, Weijie, et al.
Published: (2025)
by: He, Weijie, et al.
Published: (2025)
Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided Diffusion
by: Liang, Yijun, et al.
Published: (2024)
by: Liang, Yijun, et al.
Published: (2024)
HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Setting the Stage: Text-Driven Scene-Consistent Image Generation
by: Xie, Cong, et al.
Published: (2025)
by: Xie, Cong, et al.
Published: (2025)
DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
MotionRFT: Unified Reinforcement Fine-Tuning for Text-to-Motion Generation
by: Tan, Xiaofeng, et al.
Published: (2026)
by: Tan, Xiaofeng, et al.
Published: (2026)
SegHist: A General Segmentation-based Framework for Chinese Historical Document Text Line Detection
by: Hu, Xingjian, et al.
Published: (2024)
by: Hu, Xingjian, et al.
Published: (2024)
TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation
by: Feng, Weixi, et al.
Published: (2024)
by: Feng, Weixi, et al.
Published: (2024)
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
by: Wang, Jiarui, et al.
Published: (2025)
by: Wang, Jiarui, et al.
Published: (2025)
Similar Items
-
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
by: Wang, Shijian, et al.
Published: (2025) -
Attributed Synthetic Data Generation for Zero-shot Domain-specific Image Classification
by: Wang, Shijian, et al.
Published: (2025) -
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024) -
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
by: Wang, Shijian, et al.
Published: (2026) -
Training-free Stylized Text-to-Image Generation with Fast Inference
by: Ma, Xin, et al.
Published: (2025)