IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914068170801152 |
|---|---|
| author | Guo, Jiayi Yan, Chuanhao Xu, Xingqian Wang, Yulin Wang, Kai Huang, Gao Shi, Humphrey |
| author_facet | Guo, Jiayi Yan, Chuanhao Xu, Xingqian Wang, Yulin Wang, Kai Huang, Gao Shi, Humphrey |
| contents | Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be limited and difficult to scale up. Recent editing-based methods further refine local regions of generated images but may compromise overall image quality. In this work, we propose Implicit Multimodal Guidance (IMG), a novel re-generation-based multimodal alignment framework that requires no extra data or editing operations. Specifically, given a generated image and its prompt, IMG a) utilizes a multimodal large language model (MLLM) to identify misalignments; b) introduces an Implicit Aligner that manipulates diffusion conditioning features to reduce misalignments and enable re-generation; and c) formulates the re-alignment goal into a trainable objective, namely Iteratively Updated Preference Objective. Extensive qualitative and quantitative evaluations on SDXL, SDXL-DPO, and FLUX show that IMG outperforms existing alignment methods. Furthermore, IMG acts as a flexible plug-and-play adapter, seamlessly enhancing prior finetuning-based alignment methods. Our code will be available at https://github.com/SHI-Labs/IMG-Multimodal-Diffusion-Alignment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_26231 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance Guo, Jiayi Yan, Chuanhao Xu, Xingqian Wang, Yulin Wang, Kai Huang, Gao Shi, Humphrey Computer Vision and Pattern Recognition Ensuring precise multimodal alignment between diffusion-generated images and input prompts has been a long-standing challenge. Earlier works finetune diffusion weight using high-quality preference data, which tends to be limited and difficult to scale up. Recent editing-based methods further refine local regions of generated images but may compromise overall image quality. In this work, we propose Implicit Multimodal Guidance (IMG), a novel re-generation-based multimodal alignment framework that requires no extra data or editing operations. Specifically, given a generated image and its prompt, IMG a) utilizes a multimodal large language model (MLLM) to identify misalignments; b) introduces an Implicit Aligner that manipulates diffusion conditioning features to reduce misalignments and enable re-generation; and c) formulates the re-alignment goal into a trainable objective, namely Iteratively Updated Preference Objective. Extensive qualitative and quantitative evaluations on SDXL, SDXL-DPO, and FLUX show that IMG outperforms existing alignment methods. Furthermore, IMG acts as a flexible plug-and-play adapter, seamlessly enhancing prior finetuning-based alignment methods. Our code will be available at https://github.com/SHI-Labs/IMG-Multimodal-Diffusion-Alignment. |
| title | IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.26231 |