Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pernias, Pablo, Rampas, Dominic, Richter, Mats L., Pal, Christopher J., Aubreville, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model-based Cleaning of the QUILT-1M Pathology Dataset for Text-Conditional Image Synthesis
by: Aubreville, Marc, et al.
Published: (2024)
by: Aubreville, Marc, et al.
Published: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025)
by: Wang, Lifu, et al.
Published: (2025)
Effortless Vision-Language Model Specialization in Histopathology without Annotation
by: Qiu, Jingna, et al.
Published: (2025)
by: Qiu, Jingna, et al.
Published: (2025)
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
by: Yang, Chaojie, et al.
Published: (2026)
by: Yang, Chaojie, et al.
Published: (2026)
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
by: Gong, Chao, et al.
Published: (2024)
by: Gong, Chao, et al.
Published: (2024)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Beyond accuracy: quantifying the reliability of Multiple Instance Learning for Whole Slide Image classification
by: Keshvarikhojasteh, Hassan, et al.
Published: (2024)
by: Keshvarikhojasteh, Hassan, et al.
Published: (2024)
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
by: Tong, Shengbang, et al.
Published: (2026)
by: Tong, Shengbang, et al.
Published: (2026)
StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models
by: Zhou, Mohan, et al.
Published: (2024)
by: Zhou, Mohan, et al.
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
Decomposition Sampling for Efficient Region Annotations in Active Learning
by: Qiu, Jingna, et al.
Published: (2025)
by: Qiu, Jingna, et al.
Published: (2025)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
by: Decatur, Dale, et al.
Published: (2025)
by: Decatur, Dale, et al.
Published: (2025)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models
by: Kachlon, Eliran, et al.
Published: (2026)
by: Kachlon, Eliran, et al.
Published: (2026)
Performance evaluation of deep learning models for image analysis: considerations for visual control and statistical metrics
by: Bertram, Christof A., et al.
Published: (2026)
by: Bertram, Christof A., et al.
Published: (2026)
LSSGen: Leveraging Latent Space Scaling in Flow and Diffusion for Efficient Text to Image Generation
by: Tang, Jyun-Ze, et al.
Published: (2025)
by: Tang, Jyun-Ze, et al.
Published: (2025)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
by: Zhang, Jianyi, et al.
Published: (2024)
by: Zhang, Jianyi, et al.
Published: (2024)
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training
by: Hu, Dongting, et al.
Published: (2024)
by: Hu, Dongting, et al.
Published: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
by: Tang, Bingda, et al.
Published: (2025)
by: Tang, Bingda, et al.
Published: (2025)
Blending Concepts with Text-to-Image Diffusion Models
by: Olearo, Lorenzo, et al.
Published: (2025)
by: Olearo, Lorenzo, et al.
Published: (2025)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
by: Wang, Chenglin, et al.
Published: (2026)
by: Wang, Chenglin, et al.
Published: (2026)
Benchmarking Foundation Models for Mitotic Figure Classification
by: Ammeling, Jonas, et al.
Published: (2025)
by: Ammeling, Jonas, et al.
Published: (2025)
Domain and Content Adaptive Convolutions for Cross-Domain Adenocarcinoma Segmentation
by: Wilm, Frauke, et al.
Published: (2024)
by: Wilm, Frauke, et al.
Published: (2024)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Rethinking U-net Skip Connections for Biomedical Image Segmentation
by: Wilm, Frauke, et al.
Published: (2024)
by: Wilm, Frauke, et al.
Published: (2024)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
Efficient Text-Guided Convolutional Adapter for the Diffusion Model
by: Das, Aryan, et al.
Published: (2026)
by: Das, Aryan, et al.
Published: (2026)
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
by: Xu, Xingqian, et al.
Published: (2022)
by: Xu, Xingqian, et al.
Published: (2022)
Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023)
by: Xue, Zeyue, et al.
Published: (2023)
Is Self-Supervision Enough? Benchmarking Foundation Models Against End-to-End Training for Mitotic Figure Classification
by: Ganz, Jonathan, et al.
Published: (2024)
by: Ganz, Jonathan, et al.
Published: (2024)
Efficient Semantic Diffusion Architectures for Model Training on Synthetic Echocardiograms
by: Stojanovski, David, et al.
Published: (2024)
by: Stojanovski, David, et al.
Published: (2024)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
ECNet: Effective Controllable Text-to-Image Diffusion Models
by: Li, Sicheng, et al.
Published: (2024)
by: Li, Sicheng, et al.
Published: (2024)
Exposing Text-Image Inconsistency Using Diffusion Models
by: Huang, Mingzhen, et al.
Published: (2024)
by: Huang, Mingzhen, et al.
Published: (2024)
Segmentation-Free Guidance for Text-to-Image Diffusion Models
by: Azarian, Kambiz, et al.
Published: (2024)
by: Azarian, Kambiz, et al.
Published: (2024)
Similar Items
-
Model-based Cleaning of the QUILT-1M Pathology Dataset for Text-Conditional Image Synthesis
by: Aubreville, Marc, et al.
Published: (2024) -
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025) -
Effortless Vision-Language Model Specialization in Histopathology without Annotation
by: Qiu, Jingna, et al.
Published: (2025) -
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
by: Yang, Chaojie, et al.
Published: (2026) -
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
by: Gong, Chao, et al.
Published: (2024)