Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Sida, Huang, Siqi, Luo, Ping, Zhang, Hongyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024)
by: Huang, Sida, et al.
Published: (2024)
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025)
by: Zhu, Ruishu, et al.
Published: (2025)
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
by: Jiang, Kai, et al.
Published: (2025)
by: Jiang, Kai, et al.
Published: (2025)
Variational Positive-incentive Noise: How Noise Benefits Models
by: Zhang, Hongyuan, et al.
Published: (2023)
by: Zhang, Hongyuan, et al.
Published: (2023)
Data Augmentation of Contrastive Learning is Estimating Positive-incentive Noise
by: Zhang, Hongyuan, et al.
Published: (2024)
by: Zhang, Hongyuan, et al.
Published: (2024)
ComposeAnyone: Controllable Layout-to-Human Generation with Decoupled Multimodal Conditions
by: Zhang, Shiyue, et al.
Published: (2025)
by: Zhang, Shiyue, et al.
Published: (2025)
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
by: Zheng, Guangcong, et al.
Published: (2023)
by: Zheng, Guangcong, et al.
Published: (2023)
LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion Transformer
by: Chen, Yuzhuo, et al.
Published: (2025)
by: Chen, Yuzhuo, et al.
Published: (2025)
LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer
by: Yu, Ning, et al.
Published: (2022)
by: Yu, Ning, et al.
Published: (2022)
Layout-Guided Controllable Pathology Image Generation with In-Context Diffusion Transformers
by: Shou, Yuntao, et al.
Published: (2026)
by: Shou, Yuntao, et al.
Published: (2026)
ReLayout: Versatile and Structure-Preserving Design Layout Editing via Relation-Aware Design Reconstruction
by: Lin, Jiawei, et al.
Published: (2026)
by: Lin, Jiawei, et al.
Published: (2026)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Emu: Generative Pretraining in Multimodality
by: Sun, Quan, et al.
Published: (2023)
by: Sun, Quan, et al.
Published: (2023)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
by: Ma, Ao, et al.
Published: (2025)
by: Ma, Ao, et al.
Published: (2025)
Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model
by: Huang, Yaxuan, et al.
Published: (2025)
by: Huang, Yaxuan, et al.
Published: (2025)
CityGen: Infinite and Controllable City Layout Generation
by: Deng, Jie, et al.
Published: (2023)
by: Deng, Jie, et al.
Published: (2023)
Relation-Aware Diffusion Model for Controllable Poster Layout Generation
by: Li, Fengheng, et al.
Published: (2023)
by: Li, Fengheng, et al.
Published: (2023)
ConsistCompose: Unified Multimodal Layout Control for Image Composition
by: Shi, Xuanke, et al.
Published: (2025)
by: Shi, Xuanke, et al.
Published: (2025)
Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints
by: Fang, Chuan, et al.
Published: (2023)
by: Fang, Chuan, et al.
Published: (2023)
MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
by: Huang, Zilong, et al.
Published: (2025)
by: Huang, Zilong, et al.
Published: (2025)
STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation
by: Wang, Ruyu, et al.
Published: (2025)
by: Wang, Ruyu, et al.
Published: (2025)
ToLL: Topological Layout Learning with Asymmetric Cross-View Structural Distillation for 3D Scene Graph Generation Pretraining
by: Huang, Yucheng, et al.
Published: (2026)
by: Huang, Yucheng, et al.
Published: (2026)
Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Consistent Image Layout Editing with Diffusion Models
by: Xia, Tao, et al.
Published: (2025)
by: Xia, Tao, et al.
Published: (2025)
Layout Anything: One Transformer for Universal Room Layout Estimation
by: Mia, Md Sohag, et al.
Published: (2025)
by: Mia, Md Sohag, et al.
Published: (2025)
Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
by: Horita, Daichi, et al.
Published: (2023)
by: Horita, Daichi, et al.
Published: (2023)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
ToLo: A Two-Stage, Training-Free Layout-To-Image Generation Framework For High-Overlap Layouts
by: Huang, Linhao, et al.
Published: (2025)
by: Huang, Linhao, et al.
Published: (2025)
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
by: Cheng, Bo, et al.
Published: (2024)
by: Cheng, Bo, et al.
Published: (2024)
DogLayout: Denoising Diffusion GAN for Discrete and Continuous Layout Generation
by: Gan, Zhaoxing, et al.
Published: (2024)
by: Gan, Zhaoxing, et al.
Published: (2024)
PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Models
by: Zhang, Jinhua, et al.
Published: (2024)
by: Zhang, Jinhua, et al.
Published: (2024)
Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
by: Srivastava, Divyansh, et al.
Published: (2025)
by: Srivastava, Divyansh, et al.
Published: (2025)
Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations
by: Zhang, Yuan, et al.
Published: (2026)
by: Zhang, Yuan, et al.
Published: (2026)
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
by: Liu, Zeyang, et al.
Published: (2025)
by: Liu, Zeyang, et al.
Published: (2025)
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
by: Song, Wenhui, et al.
Published: (2025)
by: Song, Wenhui, et al.
Published: (2025)
Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM
by: Wang, Can, et al.
Published: (2024)
by: Wang, Can, et al.
Published: (2024)
Generating Synthetic Invoices via Layout-Preserving Content Replacement
by: V, Bevin, et al.
Published: (2025)
by: V, Bevin, et al.
Published: (2025)
Similar Items
-
Enhance Vision-Language Alignment with Noise
by: Huang, Sida, et al.
Published: (2024) -
ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Diffusion Models
by: Zhu, Ruishu, et al.
Published: (2025) -
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation
by: Zhang, Hui, et al.
Published: (2024) -
Explore How to Inject Beneficial Noise in MLLMs
by: Zhu, Ruishu, et al.
Published: (2025) -
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
by: Jiang, Kai, et al.
Published: (2025)