LayerComposer: Multi-Human Personalized Generation via Layered Canvas
Fuente:
arXiv
Saved in:
| Main Authors: | Qian, Guocheng Gordon, Zhang, Ruihang, Chen, Tsai-Shien, Dalva, Yusuf, Goyal, Anujraaj Argo, Menapace, Willi, Skorokhodov, Ivan, Dong, Meng, Sahni, Arpit, Ostashev, Daniil, Hu, Ju, Tulyakov, Sergey, Wang, Kuan-Chieh Jackson |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
by: Dalva, Yusuf, et al.
Published: (2025)
by: Dalva, Yusuf, et al.
Published: (2025)
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
by: Qian, Guocheng Gordon, et al.
Published: (2025)
by: Qian, Guocheng Gordon, et al.
Published: (2025)
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
by: Goyal, Anujraaj Argo, et al.
Published: (2025)
Improving Progressive Generation with Decomposable Flow Matching
by: Haji-Ali, Moayed, et al.
Published: (2025)
by: Haji-Ali, Moayed, et al.
Published: (2025)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
by: Skorokhodov, Ivan, et al.
Published: (2024)
by: Skorokhodov, Ivan, et al.
Published: (2024)
VIMI: Grounding Video Generation through Multi-modal Instruction
by: Fang, Yuwei, et al.
Published: (2024)
by: Fang, Yuwei, et al.
Published: (2024)
Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
by: Chen, Tsai-Shien, et al.
Published: (2025)
by: Chen, Tsai-Shien, et al.
Published: (2025)
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
by: Wang, Kuan-Chieh, et al.
Published: (2024)
by: Wang, Kuan-Chieh, et al.
Published: (2024)
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
by: Qian, Guocheng, et al.
Published: (2024)
by: Qian, Guocheng, et al.
Published: (2024)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
Taming Diffusion Transformer for Efficient Mobile Video Generation in Seconds
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
by: Menapace, Willi, et al.
Published: (2024)
by: Menapace, Willi, et al.
Published: (2024)
AlphaFlow: Understanding and Improving MeanFlow Models
by: Zhang, Huijie, et al.
Published: (2025)
by: Zhang, Huijie, et al.
Published: (2025)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
Multi-subject Open-set Personalization in Video Generation
by: Chen, Tsai-Shien, et al.
Published: (2025)
by: Chen, Tsai-Shien, et al.
Published: (2025)
Dynamic Concepts Personalization from Single Videos
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
EasyV2V: A High-quality Instruction-based Video Editing Framework
by: Mai, Jinjie, et al.
Published: (2025)
by: Mai, Jinjie, et al.
Published: (2025)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
Improving the Diffusability of Autoencoders
by: Skorokhodov, Ivan, et al.
Published: (2025)
by: Skorokhodov, Ivan, et al.
Published: (2025)
Mind the Time: Temporally-Controlled Multi-Event Video Generation
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
by: Patashnik, Or, et al.
Published: (2025)
by: Patashnik, Or, et al.
Published: (2025)
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
by: Haji-Ali, Moayed, et al.
Published: (2024)
by: Haji-Ali, Moayed, et al.
Published: (2024)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
by: Wu, Yushu, et al.
Published: (2025)
by: Wu, Yushu, et al.
Published: (2025)
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
by: Parihar, Rishubh, et al.
Published: (2025)
by: Parihar, Rishubh, et al.
Published: (2025)
One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
by: Haji-Ali, Moayed, et al.
Published: (2026)
by: Haji-Ali, Moayed, et al.
Published: (2026)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
Can Text-to-Video Generation help Video-Language Alignment?
by: Zanella, Luca, et al.
Published: (2025)
by: Zanella, Luca, et al.
Published: (2025)
Scaling Group Inference for Diverse and High-Quality Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
Visual Personalization Turing Test
by: Abdal, Rameen, et al.
Published: (2026)
by: Abdal, Rameen, et al.
Published: (2026)
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors
by: Dalva, Yusuf, et al.
Published: (2024)
by: Dalva, Yusuf, et al.
Published: (2024)
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026)
by: Hu, Dongting, et al.
Published: (2026)
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
by: Chen, Tsai-Shien, et al.
Published: (2024)
by: Chen, Tsai-Shien, et al.
Published: (2024)
Object-level Visual Prompts for Compositional Image Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
Promptable Game Models: Text-Guided Game Simulation via Masked Diffusion Models
by: Menapace, Willi, et al.
Published: (2023)
by: Menapace, Willi, et al.
Published: (2023)
SF-V: Single Forward Video Generation Model
by: Zhang, Zhixing, et al.
Published: (2024)
by: Zhang, Zhixing, et al.
Published: (2024)
4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
by: Mi, Zhenxing, et al.
Published: (2025)
by: Mi, Zhenxing, et al.
Published: (2025)
BodyMAP -- Jointly Predicting Body Mesh and 3D Applied Pressure Map for People in Bed
by: Tandon, Abhishek, et al.
Published: (2024)
by: Tandon, Abhishek, et al.
Published: (2024)
AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video Generation
by: Girish, Sharath, et al.
Published: (2025)
by: Girish, Sharath, et al.
Published: (2025)
Similar Items
-
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
by: Dalva, Yusuf, et al.
Published: (2025) -
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
by: Qian, Guocheng Gordon, et al.
Published: (2025) -
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
by: Goyal, Anujraaj Argo, et al.
Published: (2025) -
Improving Progressive Generation with Decomposable Flow Matching
by: Haji-Ali, Moayed, et al.
Published: (2025) -
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
by: Skorokhodov, Ivan, et al.
Published: (2024)