Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Taihang, Li, Linxuan, van de Weijer, Joost, Gao, Hongcheng, Khan, Fahad Shahbaz, Yang, Jian, Cheng, Ming-Ming, Wang, Kai, Wang, Yaxing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
by: Hu, Taihang, et al.
Published: (2025)
by: Hu, Taihang, et al.
Published: (2025)
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
Free-Lunch Color-Texture Disentanglement for Stylized Image Generation
by: Qin, Jiang, et al.
Published: (2025)
by: Qin, Jiang, et al.
Published: (2025)
Training-free image inversion for one-step diffusion models
by: Wu, Tao, et al.
Published: (2026)
by: Wu, Tao, et al.
Published: (2026)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
LocInv: Localization-aware Inversion for Text-Guided Image Editing
by: Tang, Chuanming, et al.
Published: (2024)
by: Tang, Chuanming, et al.
Published: (2024)
Leveraging Semantic Attribute Binding for Free-Lunch Color Control in Diffusion Models
by: Laria, Héctor, et al.
Published: (2025)
by: Laria, Héctor, et al.
Published: (2025)
Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic Classifier
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Covariances for Free: Exploiting Mean Distributions for Training-free Federated Learning
by: Goswami, Dipam, et al.
Published: (2024)
by: Goswami, Dipam, et al.
Published: (2024)
IterInv: Iterative Inversion for Pixel-Level T2I Models
by: Tang, Chuanming, et al.
Published: (2023)
by: Tang, Chuanming, et al.
Published: (2023)
Adversarial Concept Distillation for One-Step Diffusion Personalization
by: Yang, Yixiong, et al.
Published: (2025)
by: Yang, Yixiong, et al.
Published: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models
by: Butt, Muhammad Atif, et al.
Published: (2025)
by: Butt, Muhammad Atif, et al.
Published: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
by: Chen, Shiming, et al.
Published: (2024)
by: Chen, Shiming, et al.
Published: (2024)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Enhancing Perceptual Quality in Video Super-Resolution through Temporally-Consistent Detail Synthesis using Diffusion Models
by: Rota, Claudio, et al.
Published: (2023)
by: Rota, Claudio, et al.
Published: (2023)
Assessing Open-world Forgetting in Generative Image Model Customization
by: Laria, Héctor, et al.
Published: (2024)
by: Laria, Héctor, et al.
Published: (2024)
Diversity Has Always Been There in Your Visual Autoregressive Models
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
by: Shaker, Abdelrahman, et al.
Published: (2022)
by: Shaker, Abdelrahman, et al.
Published: (2022)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
by: Wang, Chaoyang, et al.
Published: (2024)
by: Wang, Chaoyang, et al.
Published: (2024)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
by: Dudhane, Akshay, et al.
Published: (2024)
by: Dudhane, Akshay, et al.
Published: (2024)
NumColor: Precise Numeric Color Control in Text-to-Image Generation
by: Butt, Muhammad Atif, et al.
Published: (2026)
by: Butt, Muhammad Atif, et al.
Published: (2026)
FeCAM: Exploiting the Heterogeneity of Class Distributions in Exemplar-Free Continual Learning
by: Goswami, Dipam, et al.
Published: (2023)
by: Goswami, Dipam, et al.
Published: (2023)
Addressing The Devastating Effects Of Single-Task Data Poisoning In Exemplar-Free Continual Learning
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
RainDiff: End-to-end Precipitation Nowcasting Via Token-wise Attention Diffusion
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
by: Butt, Muhammad Atif, et al.
Published: (2025)
by: Butt, Muhammad Atif, et al.
Published: (2025)
No Task Left Behind: Isotropic Model Merging with Common and Task-Specific Subspaces
by: Marczak, Daniel, et al.
Published: (2025)
by: Marczak, Daniel, et al.
Published: (2025)
ColorPeel: Color Prompt Learning with Diffusion Models via Color and Shape Disentanglement
by: Butt, Muhammad Atif, et al.
Published: (2024)
by: Butt, Muhammad Atif, et al.
Published: (2024)
AViTMP: A Tracking-Specific Transformer for Single-Branch Visual Tracking
by: Tang, Chuanming, et al.
Published: (2023)
by: Tang, Chuanming, et al.
Published: (2023)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
WaDi: Weight Direction-aware Distillation for One-step Image Synthesis
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
SeMe: Training-Free Language Model Merging via Semantic Alignment
by: Gu, Jian, et al.
Published: (2025)
by: Gu, Jian, et al.
Published: (2025)
Similar Items
-
Anchor Token Matching: Implicit Structure Locking for Training-free AR Image Editing
by: Hu, Taihang, et al.
Published: (2025) -
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
by: Li, Senmao, et al.
Published: (2023) -
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024) -
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025) -
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023)