Compositional Text-to-Image Generation with Dense Blob Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nie, Weili, Liu, Sifei, Mardani, Morteza, Liu, Chao, Eckart, Benjamin, Vahdat, Arash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
von: Feng, Weixi, et al.
Veröffentlicht: (2025)
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
von: Daras, Giannis, et al.
Veröffentlicht: (2024)
von: Daras, Giannis, et al.
Veröffentlicht: (2024)
One-step Diffusion Models with $f$-Divergence Distribution Matching
von: Xu, Yilun, et al.
Veröffentlicht: (2025)
von: Xu, Yilun, et al.
Veröffentlicht: (2025)
Transition Matching Distillation for Fast Video Generation
von: Nie, Weili, et al.
Veröffentlicht: (2026)
von: Nie, Weili, et al.
Veröffentlicht: (2026)
Fast Training of Diffusion Models with Masked Transformers
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)
DiffiT: Diffusion Vision Transformers for Image Generation
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
Truncated Consistency Models
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
von: Lee, Sangyun, et al.
Veröffentlicht: (2024)
On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
von: Liu, Chao, et al.
Veröffentlicht: (2025)
von: Liu, Chao, et al.
Veröffentlicht: (2025)
AGG: Amortized Generative 3D Gaussians for Single Image to 3D
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
von: Xu, Dejia, et al.
Veröffentlicht: (2024)
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents
von: Xu, Yilun, et al.
Veröffentlicht: (2024)
von: Xu, Yilun, et al.
Veröffentlicht: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Multi-student Diffusion Distillation for Better One-step Generators
von: Song, Yanke, et al.
Veröffentlicht: (2024)
von: Song, Yanke, et al.
Veröffentlicht: (2024)
Text-to-Image GAN with Pretrained Representations
von: You, Xiaozhou, et al.
Veröffentlicht: (2024)
von: You, Xiaozhou, et al.
Veröffentlicht: (2024)
DiffUHaul: A Training-Free Method for Object Dragging in Images
von: Avrahami, Omri, et al.
Veröffentlicht: (2024)
von: Avrahami, Omri, et al.
Veröffentlicht: (2024)
Stochastic Flow Matching for Resolving Small-Scale Physics
von: Fotiadis, Stathi, et al.
Veröffentlicht: (2024)
von: Fotiadis, Stathi, et al.
Veröffentlicht: (2024)
DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
von: Liao, Guiqiu, et al.
Veröffentlicht: (2026)
von: Liao, Guiqiu, et al.
Veröffentlicht: (2026)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
TextDestroyer: A Training- and Annotation-Free Diffusion Method for Destroying Anomal Text from Images
von: Li, Mengcheng, et al.
Veröffentlicht: (2024)
von: Li, Mengcheng, et al.
Veröffentlicht: (2024)
Mode Seeking meets Mean Seeking for Fast Long Video Generation
von: Cai, Shengqu, et al.
Veröffentlicht: (2026)
von: Cai, Shengqu, et al.
Veröffentlicht: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
von: Liu, Runtao, et al.
Veröffentlicht: (2024)
Not-So-Optimal Transport Flows for 3D Point Cloud Generation
von: Hui, Ka-Hei, et al.
Veröffentlicht: (2025)
von: Hui, Ka-Hei, et al.
Veröffentlicht: (2025)
Text-To-Image with Generative Adversarial Networks
von: Momen-Tayefeh, Mehrshad
Veröffentlicht: (2024)
von: Momen-Tayefeh, Mehrshad
Veröffentlicht: (2024)
Elucidating Optimal Reward-Diversity Tradeoffs in Text-to-Image Diffusion Models
von: Jena, Rohit, et al.
Veröffentlicht: (2024)
von: Jena, Rohit, et al.
Veröffentlicht: (2024)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
von: Li, Yi, et al.
Veröffentlicht: (2026)
von: Li, Yi, et al.
Veröffentlicht: (2026)
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
von: Clavié, Benjamin, et al.
Veröffentlicht: (2025)
von: Clavié, Benjamin, et al.
Veröffentlicht: (2025)
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
von: Zhang, Xinchen, et al.
Veröffentlicht: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
EdgeFusion: On-Device Text-to-Image Generation
von: Castells, Thibault, et al.
Veröffentlicht: (2024)
von: Castells, Thibault, et al.
Veröffentlicht: (2024)
Diffuse and Disperse: Image Generation with Representation Regularization
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
ReText: Text Boosts Generalization in Image-Based Person Re-identification
von: Mamedov, Timur, et al.
Veröffentlicht: (2026)
von: Mamedov, Timur, et al.
Veröffentlicht: (2026)
Dense Optimizer : An Information Entropy-Guided Structural Search Method for Dense-like Neural Network Design
von: Tianyuan, Liu, et al.
Veröffentlicht: (2024)
von: Tianyuan, Liu, et al.
Veröffentlicht: (2024)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
von: Wang, Ying, et al.
Veröffentlicht: (2023)
von: Wang, Ying, et al.
Veröffentlicht: (2023)
TextCraftor: Your Text Encoder Can be Image Quality Controller
von: Li, Yanyu, et al.
Veröffentlicht: (2024)
von: Li, Yanyu, et al.
Veröffentlicht: (2024)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
von: Li, Yaowei, et al.
Veröffentlicht: (2025)
von: Li, Yaowei, et al.
Veröffentlicht: (2025)
Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
von: Franchi, Gianni, et al.
Veröffentlicht: (2024)
EUGens: Efficient, Unified, and General Dense Layers
von: Kim, Sang Min, et al.
Veröffentlicht: (2024)
von: Kim, Sang Min, et al.
Veröffentlicht: (2024)
FastCLIPstyler: Optimisation-free Text-based Image Style Transfer Using Style Representations
von: Suresh, Ananda Padhmanabhan, et al.
Veröffentlicht: (2022)
von: Suresh, Ananda Padhmanabhan, et al.
Veröffentlicht: (2022)
Editing Massive Concepts in Text-to-Image Diffusion Models
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
von: Seo, Hoigi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
von: Feng, Weixi, et al.
Veröffentlicht: (2025) -
Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models
von: Daras, Giannis, et al.
Veröffentlicht: (2024) -
One-step Diffusion Models with $f$-Divergence Distribution Matching
von: Xu, Yilun, et al.
Veröffentlicht: (2025) -
Transition Matching Distillation for Fast Video Generation
von: Nie, Weili, et al.
Veröffentlicht: (2026) -
Fast Training of Diffusion Models with Masked Transformers
von: Zheng, Hongkai, et al.
Veröffentlicht: (2023)