RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Pang, Lexi, Zhang, Liheng, Ye, Hang, Ma, Xiaoxuan, Wang, Yizhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
di: Ye, Hang, et al.
Pubblicazione: (2024)
di: Ye, Hang, et al.
Pubblicazione: (2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
di: Wang, Jiawei, et al.
Pubblicazione: (2024)
Rich Human Feedback for Text-to-Image Generation
di: Liang, Youwei, et al.
Pubblicazione: (2023)
di: Liang, Youwei, et al.
Pubblicazione: (2023)
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
di: He, Tianyao, et al.
Pubblicazione: (2025)
di: He, Tianyao, et al.
Pubblicazione: (2025)
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
di: Lin, Kuan Heng, et al.
Pubblicazione: (2024)
di: Lin, Kuan Heng, et al.
Pubblicazione: (2024)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
di: Gunawan, Agus, et al.
Pubblicazione: (2025)
di: Gunawan, Agus, et al.
Pubblicazione: (2025)
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
di: Cheng, Jiaxin, et al.
Pubblicazione: (2024)
di: Cheng, Jiaxin, et al.
Pubblicazione: (2024)
Expressive Text-to-Image Generation with Rich Text
di: Ge, Songwei, et al.
Pubblicazione: (2023)
di: Ge, Songwei, et al.
Pubblicazione: (2023)
Generating Storytelling Images with Rich Chains-of-Reasoning
di: Song, Xiujie, et al.
Pubblicazione: (2025)
di: Song, Xiujie, et al.
Pubblicazione: (2025)
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
di: Lin, Jiang, et al.
Pubblicazione: (2025)
di: Lin, Jiang, et al.
Pubblicazione: (2025)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
di: Luan, Bozhi, et al.
Pubblicazione: (2024)
di: Luan, Bozhi, et al.
Pubblicazione: (2024)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
di: Ohanyan, Marianna, et al.
Pubblicazione: (2024)
di: Ohanyan, Marianna, et al.
Pubblicazione: (2024)
Inject Where It Matters: Training-Free Spatially-Adaptive Identity Preservation for Text-to-Image Personalization
di: Li, Guandong, et al.
Pubblicazione: (2026)
di: Li, Guandong, et al.
Pubblicazione: (2026)
SpotActor: Training-Free Layout-Controlled Consistent Image Generation
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
di: Wang, Jiahao, et al.
Pubblicazione: (2024)
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
di: Yu, Zhenyu, et al.
Pubblicazione: (2025)
di: Yu, Zhenyu, et al.
Pubblicazione: (2025)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
di: Jia, Mengzhao, et al.
Pubblicazione: (2024)
di: Jia, Mengzhao, et al.
Pubblicazione: (2024)
DiffArtist: Towards Structure and Appearance Controllable Image Stylization
di: Jiang, Ruixiang, et al.
Pubblicazione: (2024)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2024)
IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
di: Zeng, Bohan, et al.
Pubblicazione: (2023)
di: Zeng, Bohan, et al.
Pubblicazione: (2023)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
di: Ye, Maoyuan, et al.
Pubblicazione: (2025)
di: Ye, Maoyuan, et al.
Pubblicazione: (2025)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
di: Tang, Yolo Y., et al.
Pubblicazione: (2025)
di: Tang, Yolo Y., et al.
Pubblicazione: (2025)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
di: Liu, Biao, et al.
Pubblicazione: (2025)
di: Liu, Biao, et al.
Pubblicazione: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2023)
Visually-grounded Humanoid Agents
di: Ye, Hang, et al.
Pubblicazione: (2026)
di: Ye, Hang, et al.
Pubblicazione: (2026)
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
di: Shi, Liang, et al.
Pubblicazione: (2025)
di: Shi, Liang, et al.
Pubblicazione: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
di: Zhou, Shijie, et al.
Pubblicazione: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
di: Zhou, Dewei, et al.
Pubblicazione: (2024)
di: Zhou, Dewei, et al.
Pubblicazione: (2024)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
di: Yang, Yitong, et al.
Pubblicazione: (2024)
di: Yang, Yitong, et al.
Pubblicazione: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
di: Yang, Yue, et al.
Pubblicazione: (2025)
di: Yang, Yue, et al.
Pubblicazione: (2025)
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
di: Zhang, Rong, et al.
Pubblicazione: (2025)
di: Zhang, Rong, et al.
Pubblicazione: (2025)
SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
di: Su, Chi, et al.
Pubblicazione: (2024)
di: Su, Chi, et al.
Pubblicazione: (2024)
Risk Controlled Image Retrieval
di: Cai, Kaiwen, et al.
Pubblicazione: (2023)
di: Cai, Kaiwen, et al.
Pubblicazione: (2023)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
di: Lu, Junxin, et al.
Pubblicazione: (2026)
di: Lu, Junxin, et al.
Pubblicazione: (2026)
Sketch2Human: Deep Human Generation with Disentangled Geometry and Appearance Control
di: Qu, Linzi, et al.
Pubblicazione: (2024)
di: Qu, Linzi, et al.
Pubblicazione: (2024)
FAGStyle: Feature Augmentation on Geodesic Surface for Zero-shot Text-guided Diffusion Image Style Transfer
di: Han, Yuexing, et al.
Pubblicazione: (2024)
di: Han, Yuexing, et al.
Pubblicazione: (2024)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
di: Yao, Zebin, et al.
Pubblicazione: (2025)
di: Yao, Zebin, et al.
Pubblicazione: (2025)
A General Framework to Boost 3D GS Initialization for Text-to-3D Generation by Lexical Richness
di: Jiang, Lutao, et al.
Pubblicazione: (2024)
di: Jiang, Lutao, et al.
Pubblicazione: (2024)
Dynamic Frequency Modulation for Controllable Text-driven Image Generation
di: Shi, Tiandong, et al.
Pubblicazione: (2026)
di: Shi, Tiandong, et al.
Pubblicazione: (2026)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
di: Han, Woojung, et al.
Pubblicazione: (2025)
di: Han, Woojung, et al.
Pubblicazione: (2025)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
di: Chen, Hong, et al.
Pubblicazione: (2024)
di: Chen, Hong, et al.
Pubblicazione: (2024)
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
di: Zhang, Ruiqiang, et al.
Pubblicazione: (2026)
di: Zhang, Ruiqiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
di: Ye, Hang, et al.
Pubblicazione: (2024) -
Boximator: Generating Rich and Controllable Motions for Video Synthesis
di: Wang, Jiawei, et al.
Pubblicazione: (2024) -
Rich Human Feedback for Text-to-Image Generation
di: Liang, Youwei, et al.
Pubblicazione: (2023) -
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
di: He, Tianyao, et al.
Pubblicazione: (2025) -
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
di: Lin, Kuan Heng, et al.
Pubblicazione: (2024)