TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Jin, Xin, Zhong, Yichuan, Tian, Yapeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
di: Jin, Xin, et al.
Pubblicazione: (2025)
di: Jin, Xin, et al.
Pubblicazione: (2025)
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
di: Huang, Wenjing, et al.
Pubblicazione: (2023)
di: Huang, Wenjing, et al.
Pubblicazione: (2023)
Multi-layer Learnable Attention Mask for Multimodal Tasks
di: Barrios, Wayner, et al.
Pubblicazione: (2024)
di: Barrios, Wayner, et al.
Pubblicazione: (2024)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
di: Chen, Weifeng, et al.
Pubblicazione: (2023)
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
di: Liao, Yi, et al.
Pubblicazione: (2024)
di: Liao, Yi, et al.
Pubblicazione: (2024)
Diffusion Model-Based Video Editing: A Survey
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
di: Sun, Wenhao, et al.
Pubblicazione: (2024)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Text-to-Audio Generation Synchronized with Videos
di: Mo, Shentong, et al.
Pubblicazione: (2024)
di: Mo, Shentong, et al.
Pubblicazione: (2024)
Diffusion Models, Image Super-Resolution And Everything: A Survey
di: Moser, Brian B., et al.
Pubblicazione: (2024)
di: Moser, Brian B., et al.
Pubblicazione: (2024)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
di: Zhou, Shengli, et al.
Pubblicazione: (2026)
di: Zhou, Shengli, et al.
Pubblicazione: (2026)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
di: Wu, Xiaoping, et al.
Pubblicazione: (2024)
di: Wu, Xiaoping, et al.
Pubblicazione: (2024)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
di: Luo, Weijian, et al.
Pubblicazione: (2024)
di: Luo, Weijian, et al.
Pubblicazione: (2024)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
di: Ahire, Vrushank, et al.
Pubblicazione: (2025)
di: Ahire, Vrushank, et al.
Pubblicazione: (2025)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
di: Luo, Qiuming, et al.
Pubblicazione: (2026)
T-TAME: Trainable Attention Mechanism for Explaining Convolutional Networks and Vision Transformers
di: Ntrougkas, Mariano V., et al.
Pubblicazione: (2024)
di: Ntrougkas, Mariano V., et al.
Pubblicazione: (2024)
Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
di: Zhang, Kangkai, et al.
Pubblicazione: (2024)
Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution
di: Moser, Brian B., et al.
Pubblicazione: (2024)
di: Moser, Brian B., et al.
Pubblicazione: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
di: Zheng, Sixiao, et al.
Pubblicazione: (2025)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
Pay Less Attention to Deceptive Artifacts: Robust Detection of Compressed Deepfakes on Online Social Networks
di: Li, Manyi, et al.
Pubblicazione: (2025)
di: Li, Manyi, et al.
Pubblicazione: (2025)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
di: Mekonnen, Kidist Amde, et al.
Pubblicazione: (2024)
di: Mekonnen, Kidist Amde, et al.
Pubblicazione: (2024)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
di: Lin, Yan-Bo, et al.
Pubblicazione: (2026)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
di: Lin, Yuanze, et al.
Pubblicazione: (2025)
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
di: Zhang, Zewei, et al.
Pubblicazione: (2024)
di: Zhang, Zewei, et al.
Pubblicazione: (2024)
Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Style-Preserving Lip Sync via Audio-Aware Style Reference
di: Zhong, Weizhi, et al.
Pubblicazione: (2024)
di: Zhong, Weizhi, et al.
Pubblicazione: (2024)
Deep ReLU Networks Have Surprisingly Simple Polytopes
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
di: Barrios, Wayner, et al.
Pubblicazione: (2025)
di: Barrios, Wayner, et al.
Pubblicazione: (2025)
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
di: Kim, Ye-Chan, et al.
Pubblicazione: (2025)
di: Kim, Ye-Chan, et al.
Pubblicazione: (2025)
LayerT2V: A Unified Multi-Layer Video Generation Framework
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
di: Li, Guangzhao, et al.
Pubblicazione: (2025)
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
di: Ghosh, Dhruba, et al.
Pubblicazione: (2026)
di: Ghosh, Dhruba, et al.
Pubblicazione: (2026)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
di: Lin, Xiang, et al.
Pubblicazione: (2025)
di: Lin, Xiang, et al.
Pubblicazione: (2025)
Latent Space Probing for Adult Content Detection in Video Generative Models
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
di: Khatri, Alizishaan, et al.
Pubblicazione: (2026)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2024)
di: Liu, Sheng, et al.
Pubblicazione: (2024)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
di: Zhang, Renrui, et al.
Pubblicazione: (2023)
di: Zhang, Renrui, et al.
Pubblicazione: (2023)
Documenti analoghi
-
TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
di: Jin, Xin, et al.
Pubblicazione: (2025) -
PFB-Diff: Progressive Feature Blending Diffusion for Text-driven Image Editing
di: Huang, Wenjing, et al.
Pubblicazione: (2023) -
Multi-layer Learnable Attention Mask for Multimodal Tasks
di: Barrios, Wayner, et al.
Pubblicazione: (2024) -
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
di: Chen, Weifeng, et al.
Pubblicazione: (2023) -
Neuron Abandoning Attention Flow: Visual Explanation of Dynamics inside CNN Models
di: Liao, Yi, et al.
Pubblicazione: (2024)