Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Kuan Heng, Mo, Sicheng, Klingher, Ben, Mu, Fangzhou, Zhou, Bolei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SnAG: Scalable and Accurate Video Grounding
por: Mu, Fangzhou, et al.
Publicado: (2024)
por: Mu, Fangzhou, et al.
Publicado: (2024)
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
por: Zeng, Weichao, et al.
Publicado: (2024)
por: Zeng, Weichao, et al.
Publicado: (2024)
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
por: Wang, Hanyang, et al.
Publicado: (2026)
por: Wang, Hanyang, et al.
Publicado: (2026)
Dreamland: Controllable World Creation with Simulator and Generative Models
por: Mo, Sicheng, et al.
Publicado: (2025)
por: Mo, Sicheng, et al.
Publicado: (2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
por: He, Hao, et al.
Publicado: (2024)
por: He, Hao, et al.
Publicado: (2024)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
por: Pang, Lexi, et al.
Publicado: (2025)
por: Pang, Lexi, et al.
Publicado: (2025)
EmoCtrl: Controllable Emotional Image Content Generation
por: Yang, Jingyuan, et al.
Publicado: (2025)
por: Yang, Jingyuan, et al.
Publicado: (2025)
SimGen: Simulator-conditioned Driving Scene Generation
por: Zhou, Yunsong, et al.
Publicado: (2024)
por: Zhou, Yunsong, et al.
Publicado: (2024)
Ctrl-GenAug: Controllable Generative Augmentation for Medical Sequence Classification
por: Zhou, Xinrui, et al.
Publicado: (2024)
por: Zhou, Xinrui, et al.
Publicado: (2024)
Visual Generation Without Guidance
por: Chen, Huayu, et al.
Publicado: (2025)
por: Chen, Huayu, et al.
Publicado: (2025)
LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models
por: Butt, Muhammad Atif, et al.
Publicado: (2025)
por: Butt, Muhammad Atif, et al.
Publicado: (2025)
IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
por: Zeng, Bohan, et al.
Publicado: (2023)
por: Zeng, Bohan, et al.
Publicado: (2023)
Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints
por: Fang, Chuan, et al.
Publicado: (2023)
por: Fang, Chuan, et al.
Publicado: (2023)
DynamiCtrl: Rethinking the Basic Structure and the Role of Text for High-quality Human Image Animation
por: Zhao, Haoyu, et al.
Publicado: (2025)
por: Zhao, Haoyu, et al.
Publicado: (2025)
SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
por: Zhang, Lingyun, et al.
Publicado: (2025)
por: Zhang, Lingyun, et al.
Publicado: (2025)
SafeCtrl: Region-Aware Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
por: Zhang, Lingyun, et al.
Publicado: (2026)
por: Zhang, Lingyun, et al.
Publicado: (2026)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
por: Wang, Zhouxia, et al.
Publicado: (2023)
por: Wang, Zhouxia, et al.
Publicado: (2023)
Orchid: Image Latent Diffusion for Joint Appearance and Geometry Generation
por: Krishnan, Akshay, et al.
Publicado: (2025)
por: Krishnan, Akshay, et al.
Publicado: (2025)
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
por: Cao, Qingqing, et al.
Publicado: (2024)
por: Cao, Qingqing, et al.
Publicado: (2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
por: Lin, Han, et al.
Publicado: (2024)
por: Lin, Han, et al.
Publicado: (2024)
Steering Guidance for Personalized Text-to-Image Diffusion Models
por: Park, Sunghyun, et al.
Publicado: (2025)
por: Park, Sunghyun, et al.
Publicado: (2025)
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
por: Xu, Yijia, et al.
Publicado: (2026)
por: Xu, Yijia, et al.
Publicado: (2026)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
por: Kang, Wonjun, et al.
Publicado: (2025)
por: Kang, Wonjun, et al.
Publicado: (2025)
Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration
por: Mo, Sicheng, et al.
Publicado: (2025)
por: Mo, Sicheng, et al.
Publicado: (2025)
Street-View Image Generation from a Bird's-Eye View Layout
por: Swerdlow, Alexander, et al.
Publicado: (2023)
por: Swerdlow, Alexander, et al.
Publicado: (2023)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
Text-Driven Weakly Supervised OCT Lesion Segmentation with Structural Guidance
por: Yang, Jiaqi, et al.
Publicado: (2024)
por: Yang, Jiaqi, et al.
Publicado: (2024)
High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
por: Gao, Danyi
Publicado: (2025)
por: Gao, Danyi
Publicado: (2025)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
por: Le, Tri, et al.
Publicado: (2025)
por: Le, Tri, et al.
Publicado: (2025)
TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
por: Yang, Pei, et al.
Publicado: (2025)
por: Yang, Pei, et al.
Publicado: (2025)
Ctrl-A: Control-Driven Online Data Augmentation
por: Christensen, Jesper B., et al.
Publicado: (2026)
por: Christensen, Jesper B., et al.
Publicado: (2026)
Joint Learning of Depth and Appearance for Portrait Image Animation
por: Ji, Xinya, et al.
Publicado: (2025)
por: Ji, Xinya, et al.
Publicado: (2025)
LumiX: Structured and Coherent Text-to-Intrinsic Generation
por: Han, Xu, et al.
Publicado: (2025)
por: Han, Xu, et al.
Publicado: (2025)
EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
por: Litman, Yehonathan, et al.
Publicado: (2026)
por: Litman, Yehonathan, et al.
Publicado: (2026)
CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
por: Sun, Yiming, et al.
Publicado: (2026)
por: Sun, Yiming, et al.
Publicado: (2026)
Conditional Text-to-Image Generation with Reference Guidance
por: Kim, Taewook, et al.
Publicado: (2024)
por: Kim, Taewook, et al.
Publicado: (2024)
Estimating Appearance Models for Image Segmentation via Tensor Factorization
por: Neto, Jeova Farias Sales Rocha
Publicado: (2022)
por: Neto, Jeova Farias Sales Rocha
Publicado: (2022)
DiffArtist: Towards Structure and Appearance Controllable Image Stylization
por: Jiang, Ruixiang, et al.
Publicado: (2024)
por: Jiang, Ruixiang, et al.
Publicado: (2024)
Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models
por: Choi, Changhyun, et al.
Publicado: (2025)
por: Choi, Changhyun, et al.
Publicado: (2025)
Zero-Shot Visual Concept Blending Without Text Guidance
por: Makino, Hiroya, et al.
Publicado: (2025)
por: Makino, Hiroya, et al.
Publicado: (2025)
Ejemplares similares
-
SnAG: Scalable and Accurate Video Grounding
por: Mu, Fangzhou, et al.
Publicado: (2024) -
TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control
por: Zeng, Weichao, et al.
Publicado: (2024) -
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
por: Wang, Hanyang, et al.
Publicado: (2026) -
Dreamland: Controllable World Creation with Simulator and Generative Models
por: Mo, Sicheng, et al.
Publicado: (2025) -
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
por: He, Hao, et al.
Publicado: (2024)