SpotEdit: Selective Region Editing in Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Zhibin, Tan, Zhenxiong, Wang, Zeqing, Liu, Songhua, Wang, Xinchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
OminiControl2: Efficient Conditioning for Diffusion Transformers
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
OminiControl: Minimal and Universal Control for Diffusion Transformer
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024)
by: Wang, Shizun, et al.
Published: (2024)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
CLEAR: Conv-Like Linearization Revs Pre-Trained Diffusion Transformers Up
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
ViFeEdit: A Video-Free Tuner of Your Video Diffusion Transformer
by: Yu, Ruonan, et al.
Published: (2026)
by: Yu, Ruonan, et al.
Published: (2026)
Image Editing As Programs with Diffusion Models
by: Hu, Yujia, et al.
Published: (2025)
by: Hu, Yujia, et al.
Published: (2025)
SpotEdit: Evaluating Visually-Guided Image Editing Methods
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
by: Chen, Zigeng, et al.
Published: (2024)
by: Chen, Zigeng, et al.
Published: (2024)
Gated Condition Injection without Multimodal Attention: Towards Controllable Linear-Attention Transformers
by: Liu, Yuhe, et al.
Published: (2026)
by: Liu, Yuhe, et al.
Published: (2026)
Ultra-Resolution Adaptation with Ease
by: Yu, Ruonan, et al.
Published: (2025)
by: Yu, Ruonan, et al.
Published: (2025)
LinFusion: 1 GPU, 1 Minute, 16K Image
by: Liu, Songhua, et al.
Published: (2024)
by: Liu, Songhua, et al.
Published: (2024)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Minute-Long Videos with Dual Parallelisms
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
HP-Edit: A Human-Preference Post-Training Framework for Image Editing
by: Li, Fan, et al.
Published: (2026)
by: Li, Fan, et al.
Published: (2026)
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
by: Zhou, Letian, et al.
Published: (2025)
by: Zhou, Letian, et al.
Published: (2025)
CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
by: Xie, Weiyan, et al.
Published: (2025)
by: Xie, Weiyan, et al.
Published: (2025)
TinyFusion: Diffusion Transformers Learned Shallow
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
by: Gui, Rui, et al.
Published: (2025)
by: Gui, Rui, et al.
Published: (2025)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
by: Hui, Mude, et al.
Published: (2024)
by: Hui, Mude, et al.
Published: (2024)
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
by: Li, Zixiang, et al.
Published: (2024)
by: Li, Zixiang, et al.
Published: (2024)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
by: Shi, Zhiyi, et al.
Published: (2025)
by: Shi, Zhiyi, et al.
Published: (2025)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
by: Chen, Yuqing, et al.
Published: (2025)
by: Chen, Yuqing, et al.
Published: (2025)
GloTSFormer: Global Video Text Spotting Transformer
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
BrushEdit: All-In-One Image Inpainting and Editing
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
by: Lin, Wang, et al.
Published: (2026)
by: Lin, Wang, et al.
Published: (2026)
ExpressEdit: Fast Editing of Stylized Facial Expressions with Diffusion Models in Photoshop
by: Tang, Kenan, et al.
Published: (2026)
by: Tang, Kenan, et al.
Published: (2026)
WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
by: Kong, Hanyang, et al.
Published: (2025)
by: Kong, Hanyang, et al.
Published: (2025)
LAMS-Edit: Latent and Attention Mixing with Schedulers for Improved Content Preservation in Diffusion-Based Image and Style Editing
by: Fu, Wingwa, et al.
Published: (2026)
by: Fu, Wingwa, et al.
Published: (2026)
LatentEdit: Adaptive Latent Control for Consistent Semantic Editing
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
by: Zheng, Kaizhi, et al.
Published: (2024)
by: Zheng, Kaizhi, et al.
Published: (2024)
Region-Adaptive Sampling for Diffusion Transformers
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
Kolmogorov-Arnold Transformer
by: Yang, Xingyi, et al.
Published: (2024)
by: Yang, Xingyi, et al.
Published: (2024)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
by: Wang, Yabin, et al.
Published: (2025)
by: Wang, Yabin, et al.
Published: (2025)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
by: Jeon, Boseong, et al.
Published: (2025)
by: Jeon, Boseong, et al.
Published: (2025)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
by: Qiu, Jiaxing, et al.
Published: (2026)
by: Qiu, Jiaxing, et al.
Published: (2026)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
by: Lin, Yiqi, et al.
Published: (2026)
by: Lin, Yiqi, et al.
Published: (2026)
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
by: Tian, Feng, et al.
Published: (2024)
by: Tian, Feng, et al.
Published: (2024)
Similar Items
-
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025) -
OminiControl2: Efficient Conditioning for Diffusion Transformers
by: Tan, Zhenxiong, et al.
Published: (2025) -
OminiControl: Minimal and Universal Control for Diffusion Transformer
by: Tan, Zhenxiong, et al.
Published: (2024) -
MindBridge: A Cross-Subject Brain Decoding Framework
by: Wang, Shizun, et al.
Published: (2024) -
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)