MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Dewei, Li, You, Ma, Fan, Yang, Zongxin, Yang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
by: Zhou, Dewei, et al.
Published: (2026)
by: Zhou, Dewei, et al.
Published: (2026)
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
by: Xu, Ruihang, et al.
Published: (2025)
by: Xu, Ruihang, et al.
Published: (2025)
3DIS-FLUX: simple and efficient multi-instance generation with DiT rendering
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
by: Li, You, et al.
Published: (2026)
by: Li, You, et al.
Published: (2026)
HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
by: Zhou, Zhenglin, et al.
Published: (2024)
by: Zhou, Zhenglin, et al.
Published: (2024)
Controllable 3D Face Generation with Conditional Style Code Diffusion
by: Shen, Xiaolong, et al.
Published: (2023)
by: Shen, Xiaolong, et al.
Published: (2023)
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
by: Li, You, et al.
Published: (2024)
by: Li, You, et al.
Published: (2024)
3D Object Manipulation in a Single Image using Generative Models
by: Zhao, Ruisi, et al.
Published: (2025)
by: Zhao, Ruisi, et al.
Published: (2025)
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
by: Xu, Ruihang, et al.
Published: (2026)
by: Xu, Ruihang, et al.
Published: (2026)
Are Image-to-Video Models Good Zero-Shot Image Editors?
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy
by: Li, You, et al.
Published: (2024)
by: Li, You, et al.
Published: (2024)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction
by: Zhang, Zechuan, et al.
Published: (2023)
by: Zhang, Zechuan, et al.
Published: (2023)
GD^2-NeRF: Generative Detail Compensation via GAN and Diffusion for One-shot Generalizable Neural Radiance Fields
by: Pan, Xiao, et al.
Published: (2024)
by: Pan, Xiao, et al.
Published: (2024)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
by: Zhao, Ruisi, et al.
Published: (2026)
by: Zhao, Ruisi, et al.
Published: (2026)
MIFO: Learning and Synthesizing Multi-Instance from One Image
by: Su, Kailun, et al.
Published: (2025)
by: Su, Kailun, et al.
Published: (2025)
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
by: Hu, Jiyuan, et al.
Published: (2026)
by: Hu, Jiyuan, et al.
Published: (2026)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
InstanceV: Instance-Level Video Generation
by: Chen, Yuheng, et al.
Published: (2025)
by: Chen, Yuheng, et al.
Published: (2025)
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
by: Huang, Zehuan, et al.
Published: (2024)
by: Huang, Zehuan, et al.
Published: (2024)
ROICtrl: Boosting Instance Control for Visual Generation
by: Gu, Yuchao, et al.
Published: (2024)
by: Gu, Yuchao, et al.
Published: (2024)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
by: Yang, Zongxin, et al.
Published: (2024)
by: Yang, Zongxin, et al.
Published: (2024)
ISAC: Training-Free Instance-to-Semantic Attention Control for Improving Multi-Instance Generation
by: Jo, Sanghyun, et al.
Published: (2025)
by: Jo, Sanghyun, et al.
Published: (2025)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
by: Li, Kexin, et al.
Published: (2024)
by: Li, Kexin, et al.
Published: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation
by: Liang, Chen, et al.
Published: (2021)
by: Liang, Chen, et al.
Published: (2021)
Advancing Multiple Instance Learning with Continual Learning for Whole Slide Imaging
by: Li, Xianrui, et al.
Published: (2025)
by: Li, Xianrui, et al.
Published: (2025)
InstanceGen: Image Generation with Instance-level Instructions
by: Sella, Etai, et al.
Published: (2025)
by: Sella, Etai, et al.
Published: (2025)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
by: Xiang, Qiang, et al.
Published: (2025)
by: Xiang, Qiang, et al.
Published: (2025)
Scalable Video Object Segmentation with Identification Mechanism
by: Yang, Zongxin, et al.
Published: (2022)
by: Yang, Zongxin, et al.
Published: (2022)
Efficient Multi-Instance Generation with Janus-Pro-Dirven Prompt Parsing
by: Qi, Fan, et al.
Published: (2025)
by: Qi, Fan, et al.
Published: (2025)
Talk2Image: A Multi-Agent System for Multi-Turn Image Generation and Editing
by: Ma, Shichao, et al.
Published: (2025)
by: Ma, Shichao, et al.
Published: (2025)
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
Similar Items
-
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024) -
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
by: Zhou, Dewei, et al.
Published: (2024) -
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025) -
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
by: Zhou, Dewei, et al.
Published: (2026) -
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
by: Xu, Ruihang, et al.
Published: (2025)