Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shengyuan, Wang, Bo, Ma, Ye, Yang, Te, Cao, Xipeng, Chen, Quan, Li, Han, Dong, Di, Jiang, Peng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Text-Video Multi-Grained Integration for Video Moment Montage
by: Yin, Zhihui, et al.
Published: (2024)
by: Yin, Zhihui, et al.
Published: (2024)
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025)
by: Baek, Kanghyun, et al.
Published: (2025)
MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
by: Sala, Nathan, et al.
Published: (2026)
by: Sala, Nathan, et al.
Published: (2026)
Understanding and Improving Training-free Loss-based Diffusion Guidance
by: Shen, Yifei, et al.
Published: (2024)
by: Shen, Yifei, et al.
Published: (2024)
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)
by: Li, Jialu, et al.
Published: (2025)
Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
by: Li, Shuang, et al.
Published: (2026)
by: Li, Shuang, et al.
Published: (2026)
HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
by: Bu, Jiazi, et al.
Published: (2025)
by: Bu, Jiazi, et al.
Published: (2025)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
by: Lv, Henglei, et al.
Published: (2024)
by: Lv, Henglei, et al.
Published: (2024)
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
by: Hu, Zhenyu, et al.
Published: (2026)
by: Hu, Zhenyu, et al.
Published: (2026)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
Training-free Motion Factorization for Compositional Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
by: Deng, Yufan, et al.
Published: (2025)
by: Deng, Yufan, et al.
Published: (2025)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
by: Lee, Joohyeon, et al.
Published: (2025)
by: Lee, Joohyeon, et al.
Published: (2025)
Conditional Text-to-Image Generation with Reference Guidance
by: Kim, Taewook, et al.
Published: (2024)
by: Kim, Taewook, et al.
Published: (2024)
EEA: Exploration-Exploitation Agent for Long Video Understanding
by: Yang, Te, et al.
Published: (2025)
by: Yang, Te, et al.
Published: (2025)
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
by: Li, Yaqi, et al.
Published: (2025)
by: Li, Yaqi, et al.
Published: (2025)
Training-free Stylized Text-to-Image Generation with Fast Inference
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations
by: Li, Zejian, et al.
Published: (2024)
by: Li, Zejian, et al.
Published: (2024)
CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
by: Le, Tri, et al.
Published: (2025)
by: Le, Tri, et al.
Published: (2025)
Improving Subject-Driven Image Synthesis with Subject-Agnostic Guidance
by: Chan, Kelvin C. K., et al.
Published: (2024)
by: Chan, Kelvin C. K., et al.
Published: (2024)
Stencil: Subject-Driven Generation with Context Guidance
by: Chen, Gordon, et al.
Published: (2025)
by: Chen, Gordon, et al.
Published: (2025)
Isolated Diffusion: Optimizing Multi-Concept Text-to-Image Generation Training-Freely with Isolated Diffusion Guidance
by: Zhu, Jingyuan, et al.
Published: (2024)
by: Zhu, Jingyuan, et al.
Published: (2024)
Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects
by: Qiu, Weimin, et al.
Published: (2024)
by: Qiu, Weimin, et al.
Published: (2024)
FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
by: Xie, Qi, et al.
Published: (2025)
by: Xie, Qi, et al.
Published: (2025)
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
by: Dahary, Omer, et al.
Published: (2024)
by: Dahary, Omer, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation
by: Rahman, Aimon, et al.
Published: (2025)
by: Rahman, Aimon, et al.
Published: (2025)
Dynamic Training-Free Fusion of Subject and Style LoRAs
by: Cao, Qinglong, et al.
Published: (2026)
by: Cao, Qinglong, et al.
Published: (2026)
VideoTetris: Towards Compositional Text-to-Video Generation
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
by: Bu, Jiazi, et al.
Published: (2024)
by: Bu, Jiazi, et al.
Published: (2024)
ExpertGen: Training-Free Expert Guidance for Controllable Text-to-Face Generation
by: Shi, Liang, et al.
Published: (2025)
by: Shi, Liang, et al.
Published: (2025)
FlexiTex: Enhancing Texture Generation via Visual Guidance
by: Jiang, DaDong, et al.
Published: (2024)
by: Jiang, DaDong, et al.
Published: (2024)
FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation
by: Yao, Zebin, et al.
Published: (2025)
by: Yao, Zebin, et al.
Published: (2025)
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
by: Kou, Siqi, et al.
Published: (2026)
by: Kou, Siqi, et al.
Published: (2026)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
by: Yi, Xunpeng, et al.
Published: (2024)
by: Yi, Xunpeng, et al.
Published: (2024)
Similar Items
-
Text-Video Multi-Grained Integration for Video Moment Montage
by: Yin, Zhihui, et al.
Published: (2024) -
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
by: Baek, Kanghyun, et al.
Published: (2025) -
MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
by: Sala, Nathan, et al.
Published: (2026) -
Understanding and Improving Training-free Loss-based Diffusion Guidance
by: Shen, Yifei, et al.
Published: (2024) -
Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
by: Li, Jialu, et al.
Published: (2025)