MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Dewei, Li, You, Ma, Fan, Zhang, Xiaoting, Yang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
von: Xu, Ruihang, et al.
Veröffentlicht: (2025)
von: Xu, Ruihang, et al.
Veröffentlicht: (2025)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
von: Li, You, et al.
Veröffentlicht: (2026)
von: Li, You, et al.
Veröffentlicht: (2026)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
von: Li, You, et al.
Veröffentlicht: (2024)
von: Li, You, et al.
Veröffentlicht: (2024)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
von: Zhou, Dewei, et al.
Veröffentlicht: (2026)
von: Zhou, Dewei, et al.
Veröffentlicht: (2026)
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
von: Li, You, et al.
Veröffentlicht: (2024)
von: Li, You, et al.
Veröffentlicht: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
InstanceV: Instance-Level Video Generation
von: Chen, Yuheng, et al.
Veröffentlicht: (2025)
von: Chen, Yuheng, et al.
Veröffentlicht: (2025)
ROICtrl: Boosting Instance Control for Visual Generation
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
von: Gu, Yuchao, et al.
Veröffentlicht: (2024)
MIFO: Learning and Synthesizing Multi-Instance from One Image
von: Su, Kailun, et al.
Veröffentlicht: (2025)
von: Su, Kailun, et al.
Veröffentlicht: (2025)
Text-Guided Multi-Instance Learning for Scoliosis Screening via Gait Video Analysis
von: Li, Haiqing, et al.
Veröffentlicht: (2025)
von: Li, Haiqing, et al.
Veröffentlicht: (2025)
Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
von: Chen, Ruidong, et al.
Veröffentlicht: (2026)
von: Chen, Ruidong, et al.
Veröffentlicht: (2026)
VividDreamer: Invariant Score Distillation For Hyper-Realistic Text-to-3D Generation
von: Zhuo, Wenjie, et al.
Veröffentlicht: (2024)
von: Zhuo, Wenjie, et al.
Veröffentlicht: (2024)
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
von: Cao, Pu, et al.
Veröffentlicht: (2024)
von: Cao, Pu, et al.
Veröffentlicht: (2024)
TabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single Image
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
von: Yuan, Yu, et al.
Veröffentlicht: (2024)
ISAC: Training-Free Instance-to-Semantic Attention Control for Improving Multi-Instance Generation
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
von: Jo, Sanghyun, et al.
Veröffentlicht: (2025)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
von: Xiang, Qiang, et al.
Veröffentlicht: (2025)
von: Xiang, Qiang, et al.
Veröffentlicht: (2025)
MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
von: Tudosiu, Petru-Daniel, et al.
Veröffentlicht: (2024)
von: Tudosiu, Petru-Daniel, et al.
Veröffentlicht: (2024)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2024)
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
von: Huang, Zehuan, et al.
Veröffentlicht: (2024)
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
von: Peng, Cihang, et al.
Veröffentlicht: (2025)
InstanceGen: Image Generation with Instance-level Instructions
von: Sella, Etai, et al.
Veröffentlicht: (2025)
von: Sella, Etai, et al.
Veröffentlicht: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
von: Park, Geon, et al.
Veröffentlicht: (2025)
von: Park, Geon, et al.
Veröffentlicht: (2025)
InstanceAnimator: Multi-Instance Sketch Video Colorization
von: Zhang, Yinhan, et al.
Veröffentlicht: (2026)
von: Zhang, Yinhan, et al.
Veröffentlicht: (2026)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
von: Weng, Shuchen, et al.
Veröffentlicht: (2024)
von: Weng, Shuchen, et al.
Veröffentlicht: (2024)
MultiBooth: Towards Generating All Your Concepts in an Image from Text
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhu, Chenyang, et al.
Veröffentlicht: (2024)
Efficient Multi-Instance Generation with Janus-Pro-Dirven Prompt Parsing
von: Qi, Fan, et al.
Veröffentlicht: (2025)
von: Qi, Fan, et al.
Veröffentlicht: (2025)
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
von: Guo, Xiefan, et al.
Veröffentlicht: (2026)
von: Guo, Xiefan, et al.
Veröffentlicht: (2026)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
von: Pang, Lexi, et al.
Veröffentlicht: (2025)
von: Pang, Lexi, et al.
Veröffentlicht: (2025)
M$^{3}$T2IBench: A Large-Scale Multi-Category, Multi-Instance, Multi-Relation Text-to-Image Benchmark
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024) -
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
von: Zhou, Dewei, et al.
Veröffentlicht: (2024) -
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
von: Zhou, Dewei, et al.
Veröffentlicht: (2025) -
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
von: Xu, Ruihang, et al.
Veröffentlicht: (2025) -
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
von: Li, You, et al.
Veröffentlicht: (2026)