CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Zhanxin, Zhu, Beier, Yao, Liang, Yang, Jian, Tai, Ying |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CoDi -- an exemplar-conditioned diffusion model for low-shot counting
di: Šuštar, Grega, et al.
Pubblicazione: (2025)
di: Šuštar, Grega, et al.
Pubblicazione: (2025)
CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
di: Mei, Kangfu, et al.
Pubblicazione: (2023)
di: Mei, Kangfu, et al.
Pubblicazione: (2023)
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
di: Hsin-Ying, Lee, et al.
Pubblicazione: (2025)
di: Hsin-Ying, Lee, et al.
Pubblicazione: (2025)
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
di: Chen, Zhizhou, et al.
Pubblicazione: (2026)
di: Chen, Zhizhou, et al.
Pubblicazione: (2026)
Consistent Prompting for Rehearsal-Free Continual Learning
di: Gao, Zhanxin, et al.
Pubblicazione: (2024)
di: Gao, Zhanxin, et al.
Pubblicazione: (2024)
HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
di: Guan, Shanyan, et al.
Pubblicazione: (2024)
di: Guan, Shanyan, et al.
Pubblicazione: (2024)
Geometric Disentanglement of Text Embeddings for Subject-Consistent Text-to-Image Generation using A Single Prompt
di: Li, Shangxun, et al.
Pubblicazione: (2025)
di: Li, Shangxun, et al.
Pubblicazione: (2025)
Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
di: Ma, Jian, et al.
Pubblicazione: (2023)
di: Ma, Jian, et al.
Pubblicazione: (2023)
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
di: He, Huiguo, et al.
Pubblicazione: (2024)
di: He, Huiguo, et al.
Pubblicazione: (2024)
DiP: Taming Diffusion Models in Pixel Space
di: Chen, Zhennan, et al.
Pubblicazione: (2025)
di: Chen, Zhennan, et al.
Pubblicazione: (2025)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
di: Wang, Jiajun, et al.
Pubblicazione: (2024)
di: Wang, Jiajun, et al.
Pubblicazione: (2024)
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
di: Chen, Zhennan, et al.
Pubblicazione: (2024)
di: Chen, Zhennan, et al.
Pubblicazione: (2024)
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
di: Cheng, Junhao, et al.
Pubblicazione: (2024)
di: Cheng, Junhao, et al.
Pubblicazione: (2024)
Multi Positive Contrastive Learning with Pose-Consistent Generated Images
di: Inayoshi, Sho, et al.
Pubblicazione: (2024)
di: Inayoshi, Sho, et al.
Pubblicazione: (2024)
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
di: Wang, Shulei, et al.
Pubblicazione: (2025)
di: Wang, Shulei, et al.
Pubblicazione: (2025)
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
di: Chen, Bowen, et al.
Pubblicazione: (2025)
di: Chen, Bowen, et al.
Pubblicazione: (2025)
Investigating Text Insulation and Attention Mechanisms for Complex Visual Text Generation
di: Tai, Ying, et al.
Pubblicazione: (2025)
di: Tai, Ying, et al.
Pubblicazione: (2025)
FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
di: Shi, Chuancheng, et al.
Pubblicazione: (2025)
di: Shi, Chuancheng, et al.
Pubblicazione: (2025)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
di: Ma, Yue, et al.
Pubblicazione: (2023)
di: Ma, Yue, et al.
Pubblicazione: (2023)
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
di: Liu, Yexin, et al.
Pubblicazione: (2025)
di: Liu, Yexin, et al.
Pubblicazione: (2025)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
di: Yao, Zebin, et al.
Pubblicazione: (2025)
di: Yao, Zebin, et al.
Pubblicazione: (2025)
StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
di: Gaur, Gopalji, et al.
Pubblicazione: (2025)
di: Gaur, Gopalji, et al.
Pubblicazione: (2025)
MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
di: Gao, Junyao, et al.
Pubblicazione: (2026)
di: Gao, Junyao, et al.
Pubblicazione: (2026)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
di: Chen, Chen, et al.
Pubblicazione: (2025)
di: Chen, Chen, et al.
Pubblicazione: (2025)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
di: Chen, Hong, et al.
Pubblicazione: (2023)
di: Chen, Hong, et al.
Pubblicazione: (2023)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
di: Mai, Ziyang, et al.
Pubblicazione: (2025)
di: Mai, Ziyang, et al.
Pubblicazione: (2025)
Towards Consistent Long-Term Pose Generation
di: Li, Yayuan, et al.
Pubblicazione: (2025)
di: Li, Yayuan, et al.
Pubblicazione: (2025)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
di: Yang, Mengping, et al.
Pubblicazione: (2026)
di: Yang, Mengping, et al.
Pubblicazione: (2026)
StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors
di: Sun, Xiaokun, et al.
Pubblicazione: (2024)
di: Sun, Xiaokun, et al.
Pubblicazione: (2024)
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
di: Jian, Siyong, et al.
Pubblicazione: (2026)
di: Jian, Siyong, et al.
Pubblicazione: (2026)
HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation
di: Gan, Qijun, et al.
Pubblicazione: (2025)
di: Gan, Qijun, et al.
Pubblicazione: (2025)
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
di: Ye, Tian, et al.
Pubblicazione: (2025)
di: Ye, Tian, et al.
Pubblicazione: (2025)
DiverseAR: Boosting Diversity in Bitwise Autoregressive Image Generation
di: Yang, Ying, et al.
Pubblicazione: (2025)
di: Yang, Ying, et al.
Pubblicazione: (2025)
CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
di: Zi, Bojia, et al.
Pubblicazione: (2024)
di: Zi, Bojia, et al.
Pubblicazione: (2024)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
di: Chen, Lifeng, et al.
Pubblicazione: (2025)
di: Chen, Lifeng, et al.
Pubblicazione: (2025)
DreamBarbie: Text to Barbie-Style 3D Avatars
di: Sun, Xiaokun, et al.
Pubblicazione: (2024)
di: Sun, Xiaokun, et al.
Pubblicazione: (2024)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
di: Wei, Tianyi, et al.
Pubblicazione: (2024)
di: Wei, Tianyi, et al.
Pubblicazione: (2024)
Textual Localization: Decomposing Multi-concept Images for Subject-Driven Text-to-Image Generation
di: Shentu, Junjie, et al.
Pubblicazione: (2024)
di: Shentu, Junjie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CoDi -- an exemplar-conditioned diffusion model for low-shot counting
di: Šuštar, Grega, et al.
Pubblicazione: (2025) -
CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
di: Mei, Kangfu, et al.
Pubblicazione: (2023) -
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
di: Hsin-Ying, Lee, et al.
Pubblicazione: (2025) -
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
di: Chen, Zhizhou, et al.
Pubblicazione: (2026) -
Consistent Prompting for Rehearsal-Free Continual Learning
di: Gao, Zhanxin, et al.
Pubblicazione: (2024)