Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Woojung, Lee, Yeonkyung, Kim, Chanyoung, Park, Kwanghyun, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
by: Kim, Chanyoung, et al.
Published: (2025)
by: Kim, Chanyoung, et al.
Published: (2025)
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)
by: Kim, Chanyoung, et al.
Published: (2024)
Advancing Text-Driven Chest X-Ray Generation with Policy-Based Reinforcement Learning
by: Han, Woojung, et al.
Published: (2024)
by: Han, Woojung, et al.
Published: (2024)
PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
by: Lee, Yeonkyung, et al.
Published: (2025)
by: Lee, Yeonkyung, et al.
Published: (2025)
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)
by: Kim, Chanyoung, et al.
Published: (2024)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
by: Jeong, Taejin, et al.
Published: (2026)
by: Jeong, Taejin, et al.
Published: (2026)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
by: Lee, Yeonkyung, et al.
Published: (2026)
by: Lee, Yeonkyung, et al.
Published: (2026)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
CoBra: Complementary Branch Fusing Class and Semantic Knowledge for Robust Weakly Supervised Semantic Segmentation
by: Han, Woojung, et al.
Published: (2024)
by: Han, Woojung, et al.
Published: (2024)
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Real-Time Visual Attribution Streaming in Thinking Model
by: Kang, Seil, et al.
Published: (2026)
by: Kang, Seil, et al.
Published: (2026)
SC-Pro: Training-Free Framework for Defending Unsafe Image Synthesis Attack
by: Park, Junha, et al.
Published: (2025)
by: Park, Junha, et al.
Published: (2025)
Fourier Decomposition for Explicit Representation of 3D Point Cloud Attributes
by: Kim, Donghyun, et al.
Published: (2025)
by: Kim, Donghyun, et al.
Published: (2025)
Towards Continuous Sign Language Conversation from Isolated Signs
by: Kim, Youngmin, et al.
Published: (2026)
by: Kim, Youngmin, et al.
Published: (2026)
Rethinking Graph Convolution for 2D-to-3D Hand Pose Lifting
by: Kim, Chanyoung, et al.
Published: (2026)
by: Kim, Chanyoung, et al.
Published: (2026)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
by: Lee, Joohyeon, et al.
Published: (2025)
by: Lee, Joohyeon, et al.
Published: (2025)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
by: Park, Jihun, et al.
Published: (2025)
by: Park, Jihun, et al.
Published: (2025)
WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image
by: Park, Jiwoo, et al.
Published: (2025)
by: Park, Jiwoo, et al.
Published: (2025)
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025)
by: Baek, Kanghyun, et al.
Published: (2025)
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval
by: Park, Jeong-Woo, et al.
Published: (2025)
by: Park, Jeong-Woo, et al.
Published: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
by: Hong, Sujung, et al.
Published: (2026)
by: Hong, Sujung, et al.
Published: (2026)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
DiffuseHigh: Training-free Progressive High-Resolution Image Synthesis through Structure Guidance
by: Kim, Younghyun, et al.
Published: (2024)
by: Kim, Younghyun, et al.
Published: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
by: Choi, Tae Eun, et al.
Published: (2026)
by: Choi, Tae Eun, et al.
Published: (2026)
DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image Personalization
by: Nam, Jisu, et al.
Published: (2024)
by: Nam, Jisu, et al.
Published: (2024)
Repositioning the Subject within Image
by: Wang, Yikai, et al.
Published: (2024)
by: Wang, Yikai, et al.
Published: (2024)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
by: Hwang, Geunmin, et al.
Published: (2025)
by: Hwang, Geunmin, et al.
Published: (2025)
Global Context-aware Representation Learning for Spatially Resolved Transcriptomics
by: Oh, Yunhak, et al.
Published: (2025)
by: Oh, Yunhak, et al.
Published: (2025)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model
by: Han, Seungdae, et al.
Published: (2024)
by: Han, Seungdae, et al.
Published: (2024)
Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention
by: He, Huiguo, et al.
Published: (2024)
by: He, Huiguo, et al.
Published: (2024)
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
by: Gunawan, Agus, et al.
Published: (2025)
by: Gunawan, Agus, et al.
Published: (2025)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
by: Ohanyan, Marianna, et al.
Published: (2024)
by: Ohanyan, Marianna, et al.
Published: (2024)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
Similar Items
-
Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
by: Kim, Chanyoung, et al.
Published: (2025) -
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024) -
Advancing Text-Driven Chest X-Ray Generation with Policy-Based Reinforcement Learning
by: Han, Woojung, et al.
Published: (2024) -
PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
by: Lee, Yeonkyung, et al.
Published: (2025) -
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)