Local Representative Token Guided Merging for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Min-Jeong, Kim, Hee-Dong, Lee, Seong-Whan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
by: Jung, Gunho, et al.
Published: (2025)
by: Jung, Gunho, et al.
Published: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Illuminating Salient Contributions in Neuron Activation with Attribution Equilibrium
by: Nam, Woo-Jeoung, et al.
Published: (2022)
by: Nam, Woo-Jeoung, et al.
Published: (2022)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
LUMINA-Net: Low-light Upgrade through Multi-stage Illumination and Noise Adaptation Network for Image Enhancement
by: Siddiqua, Namrah, et al.
Published: (2025)
by: Siddiqua, Namrah, et al.
Published: (2025)
Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
by: Lee, Dong Hoon, et al.
Published: (2024)
by: Lee, Dong Hoon, et al.
Published: (2024)
Token Merging for Training-Free Semantic Binding in Text-to-Image Synthesis
by: Hu, Taihang, et al.
Published: (2024)
by: Hu, Taihang, et al.
Published: (2024)
Text-Guided Variational Image Generation for Industrial Anomaly Detection and Segmentation
by: Lee, Mingyu, et al.
Published: (2024)
by: Lee, Mingyu, et al.
Published: (2024)
DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval
by: Park, Geon, et al.
Published: (2026)
by: Park, Geon, et al.
Published: (2026)
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
by: Lee, Donghwna, et al.
Published: (2024)
by: Lee, Donghwna, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
Physically Guided Visual Mass Estimation from a Single RGB Image
by: Lee, Sungjae, et al.
Published: (2026)
by: Lee, Sungjae, et al.
Published: (2026)
CW-BASS: Confidence-Weighted Boundary-Aware Learning for Semi-Supervised Semantic Segmentation
by: Tarubinga, Ebenezer, et al.
Published: (2025)
by: Tarubinga, Ebenezer, et al.
Published: (2025)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
by: Jin, Hyun-Jun, et al.
Published: (2025)
by: Jin, Hyun-Jun, et al.
Published: (2025)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
by: Lee, Uichan, et al.
Published: (2026)
by: Lee, Uichan, et al.
Published: (2026)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
TMCIR: Token Merge Benefits Composed Image Retrieval
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024)
by: Ju, Yeong-Joon, et al.
Published: (2024)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
by: Chae, Daewon, et al.
Published: (2023)
by: Chae, Daewon, et al.
Published: (2023)
Learned Image Compression and Restoration for Digital Pathology
by: Lee, SeonYeong, et al.
Published: (2025)
by: Lee, SeonYeong, et al.
Published: (2025)
mEOL: Training-Free Instruction-Guided Multimodal Embedder for Vector Graphics and Image Retrieval
by: Kim, Kyeong Seon, et al.
Published: (2026)
by: Kim, Kyeong Seon, et al.
Published: (2026)
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025)
by: Park, Jeong-Woo, et al.
Published: (2025)
Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution Guidance
by: Hong, Jung-Ho, et al.
Published: (2023)
by: Hong, Jung-Ho, et al.
Published: (2023)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
by: Min, Hyungi, et al.
Published: (2026)
by: Min, Hyungi, et al.
Published: (2026)
Image-Guided Semantic Pseudo-LiDAR Point Generation for 3D Object Detection
by: Lee, Minseung, et al.
Published: (2024)
by: Lee, Minseung, et al.
Published: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Explaining generative diffusion models via visual analysis for interpretable decision-making process
by: Park, Ji-Hoon, et al.
Published: (2024)
by: Park, Ji-Hoon, et al.
Published: (2024)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
by: Lee, Dong-Jae, et al.
Published: (2026)
by: Lee, Dong-Jae, et al.
Published: (2026)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
by: Kim, Jihyo, et al.
Published: (2024)
by: Kim, Jihyo, et al.
Published: (2024)
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
by: Lee, Wonkwang, et al.
Published: (2025)
by: Lee, Wonkwang, et al.
Published: (2025)
Zero-shot Text-guided Infinite Image Synthesis with LLM guidance
by: Kwon, Soyeong, et al.
Published: (2024)
by: Kwon, Soyeong, et al.
Published: (2024)
Similar Items
-
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025) -
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025) -
Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
by: Jung, Gunho, et al.
Published: (2025) -
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025) -
Illuminating Salient Contributions in Neuron Activation with Attribution Equilibrium
by: Nam, Woo-Jeoung, et al.
Published: (2022)