PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yibin, Zhang, Weizhong, Zheng, Jianwei, Jin, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
DreamText: High Fidelity Scene Text Synthesis
by: Wang, Yibin, et al.
Published: (2024)
by: Wang, Yibin, et al.
Published: (2024)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
by: Zhang, Hongxiang, et al.
Published: (2024)
by: Zhang, Hongxiang, et al.
Published: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
TMCIR: Token Merge Benefits Composed Image Retrieval
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Dual Relation Alignment for Composed Image Retrieval
by: Jiang, Xintong, et al.
Published: (2023)
by: Jiang, Xintong, et al.
Published: (2023)
Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation
by: Ye, Zilyu, et al.
Published: (2024)
by: Ye, Zilyu, et al.
Published: (2024)
Enhancing Object Coherence in Layout-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
AID: Attention Interpolation of Text-to-Image Diffusion
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
by: Ding, Wei, et al.
Published: (2026)
by: Ding, Wei, et al.
Published: (2026)
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
by: Zhao, Chenchen, et al.
Published: (2026)
by: Zhao, Chenchen, et al.
Published: (2026)
UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
by: Wang, Haoxuan, et al.
Published: (2025)
by: Wang, Haoxuan, et al.
Published: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
Inline Critic Steers Image Editing
by: Kang, Weitai, et al.
Published: (2026)
by: Kang, Weitai, et al.
Published: (2026)
Decoding Vision Transformers: the Diffusion Steering Lens
by: Takatsuki, Ryota, et al.
Published: (2025)
by: Takatsuki, Ryota, et al.
Published: (2025)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives
by: Feng, Zhangchi, et al.
Published: (2024)
by: Feng, Zhangchi, et al.
Published: (2024)
A Framework For Image Synthesis Using Supervised Contrastive Learning
by: Liu, Yibin, et al.
Published: (2024)
by: Liu, Yibin, et al.
Published: (2024)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
by: Kim, Sungnyun, et al.
Published: (2023)
by: Kim, Sungnyun, et al.
Published: (2023)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
by: Zou, Siyu, et al.
Published: (2024)
by: Zou, Siyu, et al.
Published: (2024)
Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
Causally Steered Diffusion for Automated Video Counterfactual Generation
by: Spyrou, Nikos, et al.
Published: (2025)
by: Spyrou, Nikos, et al.
Published: (2025)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
by: Liu, Mingyu, et al.
Published: (2026)
by: Liu, Mingyu, et al.
Published: (2026)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
by: Zhang, Hongyu, et al.
Published: (2025)
by: Zhang, Hongyu, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
by: Bao, Zhipeng, et al.
Published: (2023)
by: Bao, Zhipeng, et al.
Published: (2023)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
by: Huang, Nisha, et al.
Published: (2024)
by: Huang, Nisha, et al.
Published: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Composing Concepts from Images and Videos via Concept-prompt Binding
by: Kong, Xianghao, et al.
Published: (2025)
by: Kong, Xianghao, et al.
Published: (2025)
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024)
by: Yamada, Moyuru
Published: (2024)
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
Multi-view Image Diffusion via Coordinate Noise and Fourier Attention
by: Theiss, Justin, et al.
Published: (2024)
by: Theiss, Justin, et al.
Published: (2024)
Paired Image Generation with Diffusion-Guided Diffusion Models
by: Zhang, Haoxuan, et al.
Published: (2025)
by: Zhang, Haoxuan, et al.
Published: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
by: Zhou, Yinan, et al.
Published: (2025)
by: Zhou, Yinan, et al.
Published: (2025)
Re-Attentional Controllable Video Diffusion Editing
by: Wang, Yuanzhi, et al.
Published: (2024)
by: Wang, Yuanzhi, et al.
Published: (2024)
Towards Accurate and Interpretable Neuroblastoma Diagnosis via Contrastive Multi-scale Pathological Image Analysis
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
by: Mei, Kangfu, et al.
Published: (2023)
by: Mei, Kangfu, et al.
Published: (2023)
Similar Items
-
High-fidelity Person-centric Subject-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023) -
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
by: Wang, Yibin, et al.
Published: (2024) -
DreamText: High Fidelity Scene Text Synthesis
by: Wang, Yibin, et al.
Published: (2024) -
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025) -
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
by: Zhang, Hongxiang, et al.
Published: (2024)