RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Lexi, Zhang, Liheng, Ye, Hang, Ma, Xiaoxuan, Wang, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
by: Ye, Hang, et al.
Published: (2024)
by: Ye, Hang, et al.
Published: (2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Rich Human Feedback for Text-to-Image Generation
by: Liang, Youwei, et al.
Published: (2023)
by: Liang, Youwei, et al.
Published: (2023)
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
by: He, Tianyao, et al.
Published: (2025)
by: He, Tianyao, et al.
Published: (2025)
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024)
by: Lin, Kuan Heng, et al.
Published: (2024)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
by: Cheng, Jiaxin, et al.
Published: (2024)
by: Cheng, Jiaxin, et al.
Published: (2024)
Expressive Text-to-Image Generation with Rich Text
by: Ge, Songwei, et al.
Published: (2023)
by: Ge, Songwei, et al.
Published: (2023)
Generating Storytelling Images with Rich Chains-of-Reasoning
by: Song, Xiujie, et al.
Published: (2025)
by: Song, Xiujie, et al.
Published: (2025)
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
by: Lin, Jiang, et al.
Published: (2025)
by: Lin, Jiang, et al.
Published: (2025)
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
by: Luan, Bozhi, et al.
Published: (2024)
by: Luan, Bozhi, et al.
Published: (2024)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
by: Ohanyan, Marianna, et al.
Published: (2024)
by: Ohanyan, Marianna, et al.
Published: (2024)
Inject Where It Matters: Training-Free Spatially-Adaptive Identity Preservation for Text-to-Image Personalization
by: Li, Guandong, et al.
Published: (2026)
by: Li, Guandong, et al.
Published: (2026)
SpotActor: Training-Free Layout-Controlled Consistent Image Generation
by: Wang, Jiahao, et al.
Published: (2024)
by: Wang, Jiahao, et al.
Published: (2024)
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
by: Jia, Mengzhao, et al.
Published: (2024)
by: Jia, Mengzhao, et al.
Published: (2024)
DiffArtist: Towards Structure and Appearance Controllable Image Stylization
by: Jiang, Ruixiang, et al.
Published: (2024)
by: Jiang, Ruixiang, et al.
Published: (2024)
IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
by: Zeng, Bohan, et al.
Published: (2023)
by: Zeng, Bohan, et al.
Published: (2023)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
by: Zhang, Yanzhe, et al.
Published: (2023)
by: Zhang, Yanzhe, et al.
Published: (2023)
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026)
by: Ye, Hang, et al.
Published: (2026)
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
by: Shi, Liang, et al.
Published: (2025)
by: Shi, Liang, et al.
Published: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
by: Yang, Yitong, et al.
Published: (2024)
by: Yang, Yitong, et al.
Published: (2024)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
FashionMAC: Deformation-Free Fashion Image Generation with Fine-Grained Model Appearance Customization
by: Zhang, Rong, et al.
Published: (2025)
by: Zhang, Rong, et al.
Published: (2025)
SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
by: Su, Chi, et al.
Published: (2024)
by: Su, Chi, et al.
Published: (2024)
Risk Controlled Image Retrieval
by: Cai, Kaiwen, et al.
Published: (2023)
by: Cai, Kaiwen, et al.
Published: (2023)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
Sketch2Human: Deep Human Generation with Disentangled Geometry and Appearance Control
by: Qu, Linzi, et al.
Published: (2024)
by: Qu, Linzi, et al.
Published: (2024)
FAGStyle: Feature Augmentation on Geodesic Surface for Zero-shot Text-guided Diffusion Image Style Transfer
by: Han, Yuexing, et al.
Published: (2024)
by: Han, Yuexing, et al.
Published: (2024)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
by: Yao, Zebin, et al.
Published: (2025)
by: Yao, Zebin, et al.
Published: (2025)
A General Framework to Boost 3D GS Initialization for Text-to-3D Generation by Lexical Richness
by: Jiang, Lutao, et al.
Published: (2024)
by: Jiang, Lutao, et al.
Published: (2024)
Dynamic Frequency Modulation for Controllable Text-driven Image Generation
by: Shi, Tiandong, et al.
Published: (2026)
by: Shi, Tiandong, et al.
Published: (2026)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
by: Han, Woojung, et al.
Published: (2025)
by: Han, Woojung, et al.
Published: (2025)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control
by: Chen, Hong, et al.
Published: (2024)
by: Chen, Hong, et al.
Published: (2024)
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
by: Zhang, Ruiqiang, et al.
Published: (2026)
by: Zhang, Ruiqiang, et al.
Published: (2026)
Similar Items
-
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
by: Ye, Hang, et al.
Published: (2024) -
Boximator: Generating Rich and Controllable Motions for Video Synthesis
by: Wang, Jiawei, et al.
Published: (2024) -
Rich Human Feedback for Text-to-Image Generation
by: Liang, Youwei, et al.
Published: (2023) -
Flux-Sculptor: Text-Driven Rich-Attribute Portrait Editing through Decomposed Spatial Flow Control
by: He, Tianyao, et al.
Published: (2025) -
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024)