Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ruidong, Bai, Yancheng, Zhang, Xuanpu, Zeng, Jianhao, Wang, Lanjun, Song, Dan, Sun, Lei, Chu, Xiangxiang, Liu, Anan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
by: Zeng, Jianhao, et al.
Published: (2025)
by: Zeng, Jianhao, et al.
Published: (2025)
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
by: He, Haodong, et al.
Published: (2025)
by: He, Haodong, et al.
Published: (2025)
SCALAR: Scale-wise Controllable Visual Autoregressive Learning
by: Xu, Ryan, et al.
Published: (2025)
by: Xu, Ryan, et al.
Published: (2025)
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
by: He, Haodong, et al.
Published: (2026)
by: He, Haodong, et al.
Published: (2026)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Semantic Context Matters: Improving Conditioning for Autoregressive Models
by: Jin, Dongyang, et al.
Published: (2025)
by: Jin, Dongyang, et al.
Published: (2025)
FLUX-Text: A Simple and Advanced Diffusion Transformer Baseline for Scene Text Editing
by: Lan, Rui, et al.
Published: (2025)
by: Lan, Rui, et al.
Published: (2025)
Group Relative Attention Guidance for Image Editing
by: Zhang, Xuanpu, et al.
Published: (2025)
by: Zhang, Xuanpu, et al.
Published: (2025)
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
by: Yu, Meng, et al.
Published: (2026)
by: Yu, Meng, et al.
Published: (2026)
CAT-DM: Controllable Accelerated Virtual Try-on with Diffusion Model
by: Zeng, Jianhao, et al.
Published: (2023)
by: Zeng, Jianhao, et al.
Published: (2023)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
by: Li, Mingxing, et al.
Published: (2025)
by: Li, Mingxing, et al.
Published: (2025)
TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models
by: Chen, Ruidong, et al.
Published: (2025)
by: Chen, Ruidong, et al.
Published: (2025)
BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data Training
by: Zhang, Xuanpu, et al.
Published: (2024)
by: Zhang, Xuanpu, et al.
Published: (2024)
Revealing Vulnerabilities in Stable Diffusion via Targeted Attacks
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
Better Fit: Accommodate Variations in Clothing Types for Virtual Try-on
by: Song, Dan, et al.
Published: (2024)
by: Song, Dan, et al.
Published: (2024)
Towards Deconfounded Image-Text Matching with Causal Inference
by: Li, Wenhui, et al.
Published: (2024)
by: Li, Wenhui, et al.
Published: (2024)
Adversarial Attacks and Defenses on Text-to-Image Diffusion Models: A Survey
by: Zhang, Chenyu, et al.
Published: (2024)
by: Zhang, Chenyu, et al.
Published: (2024)
Learning to Transform for Generalizable Instance-wise Invariance
by: Singhal, Utkarsh, et al.
Published: (2023)
by: Singhal, Utkarsh, et al.
Published: (2023)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Q-Hawkeye: Reliable Visual Policy Optimization for Image Quality Assessment
by: Xie, Wulin, et al.
Published: (2026)
by: Xie, Wulin, et al.
Published: (2026)
From Scale to Speed: Adaptive Test-Time Scaling for Image Editing
by: Qu, Xiangyan, et al.
Published: (2026)
by: Qu, Xiangyan, et al.
Published: (2026)
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
by: Fang, Xueji, et al.
Published: (2026)
by: Fang, Xueji, et al.
Published: (2026)
Image-Based Virtual Try-On: A Survey
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
Occlusion-Ordered Semantic Instance Segmentation
by: Baselizadeh, Soroosh, et al.
Published: (2025)
by: Baselizadeh, Soroosh, et al.
Published: (2025)
Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification
by: Song, Dan, et al.
Published: (2025)
by: Song, Dan, et al.
Published: (2025)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025)
by: Lu, Runnan, et al.
Published: (2025)
Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers
by: Zhang, Chunyang, et al.
Published: (2025)
by: Zhang, Chunyang, et al.
Published: (2025)
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Layer-wise Derivative Controlled Networks
by: Martnishn, Rowan, et al.
Published: (2026)
by: Martnishn, Rowan, et al.
Published: (2026)
Instance-level Visual Active Tracking with Occlusion-Aware Planning
by: Sun, Haowei, et al.
Published: (2026)
by: Sun, Haowei, et al.
Published: (2026)
T2IW: Joint Text to Image & Watermark Generation
by: Liu, An-An, et al.
Published: (2023)
by: Liu, An-An, et al.
Published: (2023)
Metaphor-based Jailbreak Attacks on Text-to-Image Models
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
by: Wu, You, et al.
Published: (2025)
by: Wu, You, et al.
Published: (2025)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
by: Qiu, Zeju, et al.
Published: (2023)
by: Qiu, Zeju, et al.
Published: (2023)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
by: Zhou, Dewei, et al.
Published: (2024)
by: Zhou, Dewei, et al.
Published: (2024)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
by: Tang, Zhicong, et al.
Published: (2026)
by: Tang, Zhicong, et al.
Published: (2026)
DanceText: A Training-Free Layered Framework for Controllable Multilingual Text Transformation in Images
by: Yu, Zhenyu, et al.
Published: (2025)
by: Yu, Zhenyu, et al.
Published: (2025)
Similar Items
-
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
by: Zeng, Jianhao, et al.
Published: (2025) -
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
by: He, Haodong, et al.
Published: (2025) -
SCALAR: Scale-wise Controllable Visual Autoregressive Learning
by: Xu, Ryan, et al.
Published: (2025) -
TEXTS-Diff: TEXTS-Aware Diffusion Model for Real-World Text Image Super-Resolution
by: He, Haodong, et al.
Published: (2026) -
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
by: Zhang, Chenyu, et al.
Published: (2025)