PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yibin, Zhang, Weizhong, Zheng, Jianwei, Jin, Cheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
High-fidelity Person-centric Subject-to-Image Synthesis
por: Wang, Yibin, et al.
Publicado: (2023)
por: Wang, Yibin, et al.
Publicado: (2023)
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
por: Wang, Yibin, et al.
Publicado: (2024)
por: Wang, Yibin, et al.
Publicado: (2024)
DreamText: High Fidelity Scene Text Synthesis
por: Wang, Yibin, et al.
Publicado: (2024)
por: Wang, Yibin, et al.
Publicado: (2024)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
por: Chang, Cheng-Hong, et al.
Publicado: (2025)
por: Chang, Cheng-Hong, et al.
Publicado: (2025)
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
por: Zhang, Hongxiang, et al.
Publicado: (2024)
por: Zhang, Hongxiang, et al.
Publicado: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
por: Jin, Zhe, et al.
Publicado: (2025)
por: Jin, Zhe, et al.
Publicado: (2025)
TMCIR: Token Merge Benefits Composed Image Retrieval
por: Wang, Chaoyang, et al.
Publicado: (2025)
por: Wang, Chaoyang, et al.
Publicado: (2025)
Dual Relation Alignment for Composed Image Retrieval
por: Jiang, Xintong, et al.
Publicado: (2023)
por: Jiang, Xintong, et al.
Publicado: (2023)
Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation
por: Ye, Zilyu, et al.
Publicado: (2024)
por: Ye, Zilyu, et al.
Publicado: (2024)
Enhancing Object Coherence in Layout-to-Image Synthesis
por: Wang, Yibin, et al.
Publicado: (2023)
por: Wang, Yibin, et al.
Publicado: (2023)
AID: Attention Interpolation of Text-to-Image Diffusion
por: He, Qiyuan, et al.
Publicado: (2024)
por: He, Qiyuan, et al.
Publicado: (2024)
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
por: Ding, Wei, et al.
Publicado: (2026)
por: Ding, Wei, et al.
Publicado: (2026)
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
por: Zhao, Chenchen, et al.
Publicado: (2026)
por: Zhao, Chenchen, et al.
Publicado: (2026)
UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer
por: Wang, Haoxuan, et al.
Publicado: (2025)
por: Wang, Haoxuan, et al.
Publicado: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
por: Taghipour, Ashkan, et al.
Publicado: (2024)
por: Taghipour, Ashkan, et al.
Publicado: (2024)
Inline Critic Steers Image Editing
por: Kang, Weitai, et al.
Publicado: (2026)
por: Kang, Weitai, et al.
Publicado: (2026)
Decoding Vision Transformers: the Diffusion Steering Lens
por: Takatsuki, Ryota, et al.
Publicado: (2025)
por: Takatsuki, Ryota, et al.
Publicado: (2025)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives
por: Feng, Zhangchi, et al.
Publicado: (2024)
por: Feng, Zhangchi, et al.
Publicado: (2024)
A Framework For Image Synthesis Using Supervised Contrastive Learning
por: Liu, Yibin, et al.
Publicado: (2024)
por: Liu, Yibin, et al.
Publicado: (2024)
DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
por: Kim, Sungnyun, et al.
Publicado: (2023)
por: Kim, Sungnyun, et al.
Publicado: (2023)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
por: Zou, Siyu, et al.
Publicado: (2024)
por: Zou, Siyu, et al.
Publicado: (2024)
Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
por: Li, Jun, et al.
Publicado: (2025)
por: Li, Jun, et al.
Publicado: (2025)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
por: Yuan, Peng, et al.
Publicado: (2026)
por: Yuan, Peng, et al.
Publicado: (2026)
Causally Steered Diffusion for Automated Video Counterfactual Generation
por: Spyrou, Nikos, et al.
Publicado: (2025)
por: Spyrou, Nikos, et al.
Publicado: (2025)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
por: Liu, Mingyu, et al.
Publicado: (2026)
por: Liu, Mingyu, et al.
Publicado: (2026)
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
por: Zhang, Hongyu, et al.
Publicado: (2025)
por: Zhang, Hongyu, et al.
Publicado: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
por: Liang, Cheng, et al.
Publicado: (2026)
por: Liang, Cheng, et al.
Publicado: (2026)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
Separate-and-Enhance: Compositional Finetuning for Text2Image Diffusion Models
por: Bao, Zhipeng, et al.
Publicado: (2023)
por: Bao, Zhipeng, et al.
Publicado: (2023)
CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis with Multimodal Diffusion
por: Huang, Nisha, et al.
Publicado: (2024)
por: Huang, Nisha, et al.
Publicado: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
por: Jung, Mingi, et al.
Publicado: (2025)
por: Jung, Mingi, et al.
Publicado: (2025)
Composing Concepts from Images and Videos via Concept-prompt Binding
por: Kong, Xianghao, et al.
Publicado: (2025)
por: Kong, Xianghao, et al.
Publicado: (2025)
GLoD: Composing Global Contexts and Local Details in Image Generation
por: Yamada, Moyuru
Publicado: (2024)
por: Yamada, Moyuru
Publicado: (2024)
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
por: Chang, Zhiyuan, et al.
Publicado: (2024)
por: Chang, Zhiyuan, et al.
Publicado: (2024)
Multi-view Image Diffusion via Coordinate Noise and Fourier Attention
por: Theiss, Justin, et al.
Publicado: (2024)
por: Theiss, Justin, et al.
Publicado: (2024)
Paired Image Generation with Diffusion-Guided Diffusion Models
por: Zhang, Haoxuan, et al.
Publicado: (2025)
por: Zhang, Haoxuan, et al.
Publicado: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
por: Zhou, Yinan, et al.
Publicado: (2025)
por: Zhou, Yinan, et al.
Publicado: (2025)
Re-Attentional Controllable Video Diffusion Editing
por: Wang, Yuanzhi, et al.
Publicado: (2024)
por: Wang, Yuanzhi, et al.
Publicado: (2024)
Towards Accurate and Interpretable Neuroblastoma Diagnosis via Contrastive Multi-scale Pathological Image Analysis
por: Zhu, Zhu, et al.
Publicado: (2025)
por: Zhu, Zhu, et al.
Publicado: (2025)
CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation
por: Mei, Kangfu, et al.
Publicado: (2023)
por: Mei, Kangfu, et al.
Publicado: (2023)
Ejemplares similares
-
High-fidelity Person-centric Subject-to-Image Synthesis
por: Wang, Yibin, et al.
Publicado: (2023) -
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
por: Wang, Yibin, et al.
Publicado: (2024) -
DreamText: High Fidelity Scene Text Synthesis
por: Wang, Yibin, et al.
Publicado: (2024) -
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
por: Chang, Cheng-Hong, et al.
Publicado: (2025) -
SteerDiff: Steering towards Safe Text-to-Image Diffusion Models
por: Zhang, Hongxiang, et al.
Publicado: (2024)