DiffDoctor: Diagnosing Image Diffusion Models Before Treating
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yiyang, Chen, Xi, Xu, Xiaogang, Ji, Sihui, Liu, Yu, Shen, Yujun, Zhao, Hengshuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffCamera: Arbitrary Refocusing on Images
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
FashionComposer: Compositional Fashion Image Generation
by: Ji, Sihui, et al.
Published: (2024)
by: Ji, Sihui, et al.
Published: (2024)
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
LayerFlow: A Unified Model for Layer-aware Video Generation
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
AnyDoor: Zero-shot Object-level Image Customization
by: Chen, Xi, et al.
Published: (2023)
by: Chen, Xi, et al.
Published: (2023)
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
FocalClick-XL: Towards Unified and High-quality Interactive Segmentation
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
LogoSticker: Inserting Logos into Diffusion Models for Customized Generation
by: Zhu, Mingkang, et al.
Published: (2024)
by: Zhu, Mingkang, et al.
Published: (2024)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Modular Customization of Diffusion Models via Blockwise-Parameterized Low-Rank Adaptation
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose Estimation
by: Xu, Li, et al.
Published: (2023)
by: Xu, Li, et al.
Published: (2023)
Seg-VAR: Image Segmentation with Visual Autoregressive Modeling
by: Zheng, Rongkun, et al.
Published: (2025)
by: Zheng, Rongkun, et al.
Published: (2025)
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
DiffGS: Functional Gaussian Splatting Diffusion
by: Zhou, Junsheng, et al.
Published: (2024)
by: Zhou, Junsheng, et al.
Published: (2024)
MacDiff: Unified Skeleton Modeling with Masked Conditional Diffusion
by: Wu, Lehong, et al.
Published: (2024)
by: Wu, Lehong, et al.
Published: (2024)
FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring
by: Liu, Xiaoyang, et al.
Published: (2025)
by: Liu, Xiaoyang, et al.
Published: (2025)
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
by: Feng, Yutong, et al.
Published: (2023)
by: Feng, Yutong, et al.
Published: (2023)
Diffusion Noise Feature: Accurate and Fast Generated Image Detection
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
FairDiff: Fair Segmentation with Point-Image Diffusion
by: Li, Wenyi, et al.
Published: (2024)
by: Li, Wenyi, et al.
Published: (2024)
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
by: Zou, Shilong, et al.
Published: (2025)
by: Zou, Shilong, et al.
Published: (2025)
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
AsyncDiff: Asynchronous Timestep Conditioning for Enhanced Text-to-Image Diffusion Inference
by: Xu, Longhuan, et al.
Published: (2025)
by: Xu, Longhuan, et al.
Published: (2025)
Depth Anything V2
by: Yang, Lihe, et al.
Published: (2024)
by: Yang, Lihe, et al.
Published: (2024)
Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning
by: Chen, Yiyang, et al.
Published: (2025)
by: Chen, Yiyang, et al.
Published: (2025)
CRS-Diff: Controllable Remote Sensing Image Generation with Diffusion Model
by: Tang, Datao, et al.
Published: (2024)
by: Tang, Datao, et al.
Published: (2024)
Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection
by: Ding, Kaixin, et al.
Published: (2025)
by: Ding, Kaixin, et al.
Published: (2025)
DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior
by: Lin, Xinqi, et al.
Published: (2023)
by: Lin, Xinqi, et al.
Published: (2023)
DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion
by: Li, Haoteng, et al.
Published: (2025)
by: Li, Haoteng, et al.
Published: (2025)
ResDiff: Combining CNN and Diffusion Model for Image Super-Resolution
by: Shang, Shuyao, et al.
Published: (2023)
by: Shang, Shuyao, et al.
Published: (2023)
BlindDiff: Empowering Degradation Modelling in Diffusion Models for Blind Image Super-Resolution
by: Li, Feng, et al.
Published: (2024)
by: Li, Feng, et al.
Published: (2024)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
Similar Items
-
DiffCamera: Arbitrary Refocusing on Images
by: Wang, Yiyang, et al.
Published: (2025) -
FashionComposer: Compositional Fashion Image Generation
by: Ji, Sihui, et al.
Published: (2024) -
GDRO: Group-level Reward Post-training Suitable for Diffusion Models
by: Wang, Yiyang, et al.
Published: (2026) -
LayerFlow: A Unified Model for Layer-aware Video Generation
by: Ji, Sihui, et al.
Published: (2025) -
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)