HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Hongji, Zhou, Yucheng, Han, Wencheng, Tao, Runzhou, Qiu, Zhongying, Yang, Jianfei, Shen, Jianbing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
by: Zheng, Huan, et al.
Published: (2025)
by: Zheng, Huan, et al.
Published: (2025)
HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
High-Precision Self-Supervised Monocular Depth Estimation with Rich-Resource Prior
by: Han, Wencheng, et al.
Published: (2024)
by: Han, Wencheng, et al.
Published: (2024)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
Towards Better Cephalometric Landmark Detection with Diffusion Data Generation
by: Guo, Dongqian, et al.
Published: (2025)
by: Guo, Dongqian, et al.
Published: (2025)
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation
by: Song, Han, et al.
Published: (2026)
by: Song, Han, et al.
Published: (2026)
HiMix: Hierarchical Artifact-aware Mixup for Generalized Synthetic Image Detection
by: Zhou, Shuchang, et al.
Published: (2026)
by: Zhou, Shuchang, et al.
Published: (2026)
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
by: Li, Yaqi, et al.
Published: (2025)
by: Li, Yaqi, et al.
Published: (2025)
HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text Classification
by: Jain, Vidit, et al.
Published: (2024)
by: Jain, Vidit, et al.
Published: (2024)
Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution
by: Zheng, Huan, et al.
Published: (2024)
by: Zheng, Huan, et al.
Published: (2024)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
by: Zhou, Ziqin, et al.
Published: (2025)
by: Zhou, Ziqin, et al.
Published: (2025)
HiGen: Hierarchical Graph Generative Networks
by: Karami, Mahdi
Published: (2023)
by: Karami, Mahdi
Published: (2023)
Multimodal Large Language Models for Multi-Subject In-Context Image Generation
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
by: Yan, Tianyi, et al.
Published: (2025)
by: Yan, Tianyi, et al.
Published: (2025)
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
by: Zhou, Yucheng, et al.
Published: (2025)
by: Zhou, Yucheng, et al.
Published: (2025)
HiDiGen: Hierarchical Diffusion for B-Rep Generation with Explicit Topological Constraints
by: Liu, Shurui, et al.
Published: (2026)
by: Liu, Shurui, et al.
Published: (2026)
AdaOcc: Adaptive Forward View Transformation and Flow Modeling for 3D Occupancy and Flow Prediction
by: Chen, Dubing, et al.
Published: (2024)
by: Chen, Dubing, et al.
Published: (2024)
Towards Geometry-Aware and Motion-Guided Video Human Mesh Recovery
by: Chen, Hongjun, et al.
Published: (2026)
by: Chen, Hongjun, et al.
Published: (2026)
RAWMamba: Unified sRGB-to-RAW De-rendering With State Space Model
by: Chen, Hongjun, et al.
Published: (2024)
by: Chen, Hongjun, et al.
Published: (2024)
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
by: Cheng, Bo, et al.
Published: (2024)
by: Cheng, Bo, et al.
Published: (2024)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
HiVeGen -- Hierarchical LLM-based Verilog Generation for Scalable Chip Design
by: Tang, Jinwei, et al.
Published: (2024)
by: Tang, Jinwei, et al.
Published: (2024)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
by: Zhou, Yucheng, et al.
Published: (2025)
by: Zhou, Yucheng, et al.
Published: (2025)
Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
by: Song, Lingran, et al.
Published: (2025)
by: Song, Lingran, et al.
Published: (2025)
Probing Commonsense Reasoning Capability of Text-to-Image Generative Models via Non-visual Description
by: Pan, Mianzhi, et al.
Published: (2023)
by: Pan, Mianzhi, et al.
Published: (2023)
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
by: Han, Wencheng, et al.
Published: (2024)
by: Han, Wencheng, et al.
Published: (2024)
Diffusion Model with Representation Alignment for Protein Inverse Folding
by: Wang, Chenglin, et al.
Published: (2024)
by: Wang, Chenglin, et al.
Published: (2024)
Reducing CT Metal Artifacts by Learning Latent Space Alignment with Gemstone Spectral Imaging Data
by: Han, Wencheng, et al.
Published: (2025)
by: Han, Wencheng, et al.
Published: (2025)
Robustness of Watermarking on Text-to-Image Diffusion Models
by: Wu, Xiaodong, et al.
Published: (2024)
by: Wu, Xiaodong, et al.
Published: (2024)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
by: Wang, Chenglin, et al.
Published: (2026)
by: Wang, Chenglin, et al.
Published: (2026)
RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
by: Li, Chunliang, et al.
Published: (2024)
by: Li, Chunliang, et al.
Published: (2024)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
by: Dong, Wenqi, et al.
Published: (2025)
by: Dong, Wenqi, et al.
Published: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
GenCompositor: Generative Video Compositing with Diffusion Transformer
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
Breaking Down Monocular Ambiguity: Exploiting Temporal Evolution for 3D Lane Detection
by: Zheng, Huan, et al.
Published: (2025)
by: Zheng, Huan, et al.
Published: (2025)
Similar Items
-
Self-Rewarding Large Vision-Language Models for Optimizing Prompts in Text-to-Image Generation
by: Yang, Hongji, et al.
Published: (2025) -
DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models
by: Yang, Hongji, et al.
Published: (2025) -
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
by: Zheng, Huan, et al.
Published: (2025) -
HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
by: Yang, Hongji, et al.
Published: (2026) -
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
by: Zhou, Yucheng, et al.
Published: (2026)