LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Runhui, Cai, Kaixin, Han, Jianhua, Liang, Xiaodan, Pei, Renjing, Lu, Guansong, Xu, Songcen, Zhang, Wei, Xu, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment
by: Jiang, Zutao, et al.
Published: (2023)
by: Jiang, Zutao, et al.
Published: (2023)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023)
by: Lu, Guansong, et al.
Published: (2023)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
by: Cai, Kaixin, et al.
Published: (2026)
by: Cai, Kaixin, et al.
Published: (2026)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
by: Lu, Tianyi, et al.
Published: (2023)
by: Lu, Tianyi, et al.
Published: (2023)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025)
by: Chen, Zisheng, et al.
Published: (2025)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
by: Wang, Chunwei, et al.
Published: (2024)
by: Wang, Chunwei, et al.
Published: (2024)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
by: Shi, Fengyuan, et al.
Published: (2023)
by: Shi, Fengyuan, et al.
Published: (2023)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023)
by: Kim, Sungnyun, et al.
Published: (2023)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
by: Huang, Runhui, et al.
Published: (2025)
by: Huang, Runhui, et al.
Published: (2025)
LayeringDiff: Layered Image Synthesis via Generation, then Disassembly with Generative Knowledge
by: Kang, Kyoungkook, et al.
Published: (2025)
by: Kang, Kyoungkook, et al.
Published: (2025)
ACE: Anti-Editing Concept Erasure in Text-to-Image Models
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
by: Wang, Zitong, et al.
Published: (2025)
by: Wang, Zitong, et al.
Published: (2025)
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
by: Gu, Geonmo, et al.
Published: (2023)
by: Gu, Geonmo, et al.
Published: (2023)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024)
by: Zhang, Xing, et al.
Published: (2024)
MagicEraser: Erasing Any Objects via Semantics-Aware Control
by: Li, Fan, et al.
Published: (2024)
by: Li, Fan, et al.
Published: (2024)
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
by: Khwaja, Emaad, et al.
Published: (2024)
by: Khwaja, Emaad, et al.
Published: (2024)
Modality-Composable Diffusion Policy via Inference-Time Distribution-level Composition
by: Cao, Jiahang, et al.
Published: (2025)
by: Cao, Jiahang, et al.
Published: (2025)
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
TextDiffSeg: Text-guided Latent Diffusion Model for 3d Medical Images Segmentation
by: Ma, Kangbo
Published: (2025)
by: Ma, Kangbo
Published: (2025)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
by: Huang, Wenjing, et al.
Published: (2023)
by: Huang, Wenjing, et al.
Published: (2023)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
by: Liang, Xiwen, et al.
Published: (2023)
by: Liang, Xiwen, et al.
Published: (2023)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
by: Xiang, Kun, et al.
Published: (2024)
by: Xiang, Kun, et al.
Published: (2024)
Can Atomic Step Decomposition Enhance the Self-structured Reasoning of Multimodal Large Models?
by: Xiang, Kun, et al.
Published: (2025)
by: Xiang, Kun, et al.
Published: (2025)
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution
by: Zheng, Qingping, et al.
Published: (2024)
by: Zheng, Qingping, et al.
Published: (2024)
Phy-Diff: Physics-guided Hourglass Diffusion Model for Diffusion MRI Synthesis
by: Zhang, Juanhua, et al.
Published: (2024)
by: Zhang, Juanhua, et al.
Published: (2024)
Composing Linear Layers from Irreducibles
by: Pence, Travis, et al.
Published: (2025)
by: Pence, Travis, et al.
Published: (2025)
DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models
by: Xing, Ximing, et al.
Published: (2023)
by: Xing, Ximing, et al.
Published: (2023)
LayerComposer: Multi-Human Personalized Generation via Layered Canvas
by: Qian, Guocheng Gordon, et al.
Published: (2025)
by: Qian, Guocheng Gordon, et al.
Published: (2025)
LayerDiffusion: Layered Controlled Image Editing with Diffusion Models
by: Li, Pengzhi, et al.
Published: (2023)
by: Li, Pengzhi, et al.
Published: (2023)
LAPTOP-Diff: Layer Pruning and Normalized Distillation for Compressing Diffusion Models
by: Zhang, Dingkun, et al.
Published: (2024)
by: Zhang, Dingkun, et al.
Published: (2024)
MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation
by: Cai, Kaixin, et al.
Published: (2023)
by: Cai, Kaixin, et al.
Published: (2023)
Image-aware Layout Generation with User Constraints for Poster Design
by: Xu, Chenchen, et al.
Published: (2026)
by: Xu, Chenchen, et al.
Published: (2026)
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
by: Lin, Jingbo, et al.
Published: (2024)
by: Lin, Jingbo, et al.
Published: (2024)
Similar Items
-
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment
by: Jiang, Zutao, et al.
Published: (2023) -
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023) -
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023) -
MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
by: Cai, Kaixin, et al.
Published: (2026) -
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
by: Lu, Tianyi, et al.
Published: (2023)