Generating Compositional Scenes via Text-to-image RGBA Instance Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Fontanella, Alessandro, Tudosiu, Petru-Daniel, Yang, Yongxin, Zhang, Shifeng, Parisot, Sarah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Practical Investigation of Spatially-Controlled Image Generation with Transformers
by: Xia, Guoxuan, et al.
Published: (2025)
by: Xia, Guoxuan, et al.
Published: (2025)
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
by: Tudosiu, Petru-Daniel, et al.
Published: (2024)
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
by: Dutt, Raman, et al.
Published: (2025)
by: Dutt, Raman, et al.
Published: (2025)
SceneForge: Structured World Supervision from 3D Interventions
by: Li, Jizhizi, et al.
Published: (2026)
by: Li, Jizhizi, et al.
Published: (2026)
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
by: Liu, Minghao, et al.
Published: (2024)
by: Liu, Minghao, et al.
Published: (2024)
GALOT: Generative Active Learning via Optimizable Zero-shot Text-to-image Generation
by: Hong, Hanbin, et al.
Published: (2024)
by: Hong, Hanbin, et al.
Published: (2024)
Progressive Compositionality in Text-to-Image Generative Models
by: Han, Evans Xu, et al.
Published: (2024)
by: Han, Evans Xu, et al.
Published: (2024)
Test-Time Instance-Specific Parameter Composition: A New Paradigm for Adaptive Generative Modeling
by: Tran, Minh-Tuan, et al.
Published: (2026)
by: Tran, Minh-Tuan, et al.
Published: (2026)
TransAnimate: Taming Layer Diffusion to Generate RGBA Video
by: Chen, Xuewei, et al.
Published: (2025)
by: Chen, Xuewei, et al.
Published: (2025)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
by: Zhao, Zelin, et al.
Published: (2025)
by: Zhao, Zelin, et al.
Published: (2025)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Enhancing Compositional Generalization via Compositional Feature Alignment
by: Wang, Haoxiang, et al.
Published: (2024)
by: Wang, Haoxiang, et al.
Published: (2024)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
by: Izadi, Amir Mohammad, et al.
Published: (2025)
by: Izadi, Amir Mohammad, et al.
Published: (2025)
Laplacian Multi-scale Flow Matching for Generative Modeling
by: Zhao, Zelin, et al.
Published: (2026)
by: Zhao, Zelin, et al.
Published: (2026)
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024)
by: Fan, Lijie, et al.
Published: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
ISCUTE: Instance Segmentation of Cables Using Text Embedding
by: Kozlovsky, Shir, et al.
Published: (2024)
by: Kozlovsky, Shir, et al.
Published: (2024)
Text-to-image Diffusion Models in Generative AI: A Survey
by: Zhang, Chenshuang, et al.
Published: (2023)
by: Zhang, Chenshuang, et al.
Published: (2023)
Unbiased Scene Graph Generation from Biased Training
by: Tang, Kaihua, et al.
Published: (2020)
by: Tang, Kaihua, et al.
Published: (2020)
A Fair Ranking and New Model for Panoptic Scene Graph Generation
by: Lorenz, Julian, et al.
Published: (2024)
by: Lorenz, Julian, et al.
Published: (2024)
EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation
by: Mun, Sunung, et al.
Published: (2026)
by: Mun, Sunung, et al.
Published: (2026)
All Seeds Are Not Equal: Enhancing Compositional Text-to-Image Generation with Reliable Random Seeds
by: Li, Shuangqi, et al.
Published: (2024)
by: Li, Shuangqi, et al.
Published: (2024)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
Phrase-Instance Alignment for Generalized Referring Segmentation
by: Nguyen, E-Ro, et al.
Published: (2024)
by: Nguyen, E-Ro, et al.
Published: (2024)
gen2seg: Generative Models Enable Generalizable Instance Segmentation
by: Khangaonkar, Om, et al.
Published: (2025)
by: Khangaonkar, Om, et al.
Published: (2025)
Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
EP-Diffuser: An Efficient Diffusion Model for Traffic Scene Generation and Prediction via Polynomial Representations
by: Yao, Yue, et al.
Published: (2025)
by: Yao, Yue, et al.
Published: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Text-to-Scene with Large Reasoning Models
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
by: Zhang, Yuhang, et al.
Published: (2024)
by: Zhang, Yuhang, et al.
Published: (2024)
Decision Boundary-aware Knowledge Consolidation Generates Better Instance-Incremental Learner
by: Nie, Qiang, et al.
Published: (2024)
by: Nie, Qiang, et al.
Published: (2024)
Compositional Text-to-Image Generation with Dense Blob Representations
by: Nie, Weili, et al.
Published: (2024)
by: Nie, Weili, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2024)
by: Zhai, Guangyao, et al.
Published: (2024)
Multisource Collaborative Domain Generalization for Cross-Scene Remote Sensing Image Classification
by: Han, Zhu, et al.
Published: (2024)
by: Han, Zhu, et al.
Published: (2024)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
PILOT: A Data-Free Continual Learning Approach for Real-Time Semantic Segmentation via Boundary Guidance
by: Zhou, Yujing, et al.
Published: (2026)
by: Zhou, Yujing, et al.
Published: (2026)
Disentangled 3D Scene Generation with Layout Learning
by: Epstein, Dave, et al.
Published: (2024)
by: Epstein, Dave, et al.
Published: (2024)
Similar Items
-
A Practical Investigation of Spatially-Controlled Image Generation with Transformers
by: Xia, Guoxuan, et al.
Published: (2025) -
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
by: Tudosiu, Petru-Daniel, et al.
Published: (2024) -
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
by: Dutt, Raman, et al.
Published: (2025) -
SceneForge: Structured World Supervision from 3D Interventions
by: Li, Jizhizi, et al.
Published: (2026) -
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)