ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gal, Rinon, Haviv, Adi, Alaluf, Yuval, Bermano, Amit H., Cohen-Or, Daniel, Chechik, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LCM-Lookahead for Encoder-based Text-to-Image Personalization
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023)
by: Tewel, Yoad, et al.
Published: (2023)
ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation
by: Shalev-Arkushin, Rotem, et al.
Published: (2025)
by: Shalev-Arkushin, Rotem, et al.
Published: (2025)
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
TurboEdit: Text-Based Image Editing Using Few-Step Diffusion Models
by: Deutch, Gilad, et al.
Published: (2024)
by: Deutch, Gilad, et al.
Published: (2024)
IP-Composer: Semantic Composition of Visual Concepts
by: Dorfman, Sara, et al.
Published: (2025)
by: Dorfman, Sara, et al.
Published: (2025)
DiffUHaul: A Training-Free Method for Object Dragging in Images
by: Avrahami, Omri, et al.
Published: (2024)
by: Avrahami, Omri, et al.
Published: (2024)
TriTex: Learning Texture from a Single Mesh via Triplane Semantic Features
by: Cohen-Bar, Dana, et al.
Published: (2025)
by: Cohen-Bar, Dana, et al.
Published: (2025)
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
by: Manor, Hila, et al.
Published: (2026)
by: Manor, Hila, et al.
Published: (2026)
Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation
by: Atzmon, Yuval, et al.
Published: (2024)
by: Atzmon, Yuval, et al.
Published: (2024)
Consolidating Attention Features for Multi-view Image Editing
by: Patashnik, Or, et al.
Published: (2024)
by: Patashnik, Or, et al.
Published: (2024)
Assessing Image Quality Using a Simple Generative Representation
by: Raviv, Simon, et al.
Published: (2024)
by: Raviv, Simon, et al.
Published: (2024)
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
by: Binyamin, Lital, et al.
Published: (2024)
by: Binyamin, Lital, et al.
Published: (2024)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
by: Patashnik, Or, et al.
Published: (2025)
by: Patashnik, Or, et al.
Published: (2025)
MAS: Multi-view Ancestral Sampling for 3D motion generation using 2D diffusion
by: Kapon, Roy, et al.
Published: (2023)
by: Kapon, Roy, et al.
Published: (2023)
AnyTop: Character Animation Diffusion with Any Topology
by: Gat, Inbar, et al.
Published: (2025)
by: Gat, Inbar, et al.
Published: (2025)
PALP: Prompt Aligned Personalization of Text-to-Image Models
by: Arar, Moab, et al.
Published: (2024)
by: Arar, Moab, et al.
Published: (2024)
Policy Optimized Text-to-Image Pipeline Design
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
Padding Tone: A Mechanistic Analysis of Padding Tokens in T2I Models
by: Toker, Michael, et al.
Published: (2025)
by: Toker, Michael, et al.
Published: (2025)
Not All Similarities Are Created Equal: Leveraging Data-Driven Biases to Inform GenAI Copyright Disputes
by: Hacohen, Uri, et al.
Published: (2024)
by: Hacohen, Uri, et al.
Published: (2024)
Data-Driven Loss Functions for Inference-Time Optimization in Text-to-Image
by: Yiflach, Sapir Esther, et al.
Published: (2025)
by: Yiflach, Sapir Esther, et al.
Published: (2025)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
4-LEGS: 4D Language Embedded Gaussian Splatting
by: Fiebelman, Gal, et al.
Published: (2024)
by: Fiebelman, Gal, et al.
Published: (2024)
Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
by: Aloni, Yaron, et al.
Published: (2025)
by: Aloni, Yaron, et al.
Published: (2025)
Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
by: Barda, Amir, et al.
Published: (2024)
by: Barda, Amir, et al.
Published: (2024)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
by: Almog, Gal, et al.
Published: (2025)
by: Almog, Gal, et al.
Published: (2025)
Monkey See, Monkey Do: Harnessing Self-attention in Motion Diffusion for Zero-shot Motion Transfer
by: Raab, Sigal, et al.
Published: (2024)
by: Raab, Sigal, et al.
Published: (2024)
CAP: Evaluation of Persuasive and Creative Image Generation
by: Aghazadeh, Aysan, et al.
Published: (2024)
by: Aghazadeh, Aysan, et al.
Published: (2024)
Masked Extended Attention for Zero-Shot Virtual Try-On In The Wild
by: Orzech, Nadav, et al.
Published: (2024)
by: Orzech, Nadav, et al.
Published: (2024)
Random Walks in Self-supervised Learning for Triangular Meshes
by: Yefet, Gal, et al.
Published: (2025)
by: Yefet, Gal, et al.
Published: (2025)
Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
by: Fiebelman, Gal, et al.
Published: (2025)
by: Fiebelman, Gal, et al.
Published: (2025)
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
by: Zafar, Oz, et al.
Published: (2024)
by: Zafar, Oz, et al.
Published: (2024)
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
by: S, Sridhar, et al.
Published: (2025)
by: S, Sridhar, et al.
Published: (2025)
JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion
by: Chen, Anthony, et al.
Published: (2026)
by: Chen, Anthony, et al.
Published: (2026)
MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
by: Sala, Nathan, et al.
Published: (2026)
by: Sala, Nathan, et al.
Published: (2026)
Spice-E : Structural Priors in 3D Diffusion using Cross-Entity Attention
by: Sella, Etai, et al.
Published: (2023)
by: Sella, Etai, et al.
Published: (2023)
Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes
by: Krakovsky, Shai, et al.
Published: (2025)
by: Krakovsky, Shai, et al.
Published: (2025)
MG-Gen: Single Image to Motion Graphics Generation
by: Shirakawa, Takahiro, et al.
Published: (2025)
by: Shirakawa, Takahiro, et al.
Published: (2025)
Object-level Visual Prompts for Compositional Image Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
Similar Items
-
LCM-Lookahead for Encoder-based Text-to-Image Personalization
by: Gal, Rinon, et al.
Published: (2024) -
Key-Locked Rank One Editing for Text-to-Image Personalization
by: Tewel, Yoad, et al.
Published: (2023) -
ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation
by: Shalev-Arkushin, Rotem, et al.
Published: (2025) -
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024) -
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)