The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908675966238720 |
|---|---|
| author | Ouyang, Ziheng Song, Yiren Liu, Yaoli Zhu, Shihao Hou, Qibin Cheng, Ming-Ming Shou, Mike Zheng |
| author_facet | Ouyang, Ziheng Song, Yiren Liu, Yaoli Zhu, Shihao Hou, Qibin Cheng, Ming-Ming Shou, Mike Zheng |
| contents | Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approach and present our ImageCritic. We first construct a dataset of reference-degraded-target triplets obtained via VLM-based selection and explicit degradation, which effectively simulates the common inaccuracies or inconsistencies observed in existing generation models. Furthermore, building on a thorough examination of the model's attention mechanisms and intrinsic representations, we accordingly devise an attention alignment loss and a detail encoder to precisely rectify inconsistencies. ImageCritic can be integrated into an agent framework to automatically detect inconsistencies and correct them with multi-round and local editing in complex scenarios. Extensive experiments demonstrate that ImageCritic can effectively resolve detail-related issues in various customized generation scenarios, providing significant improvements over existing methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_20614 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment Ouyang, Ziheng Song, Yiren Liu, Yaoli Zhu, Shihao Hou, Qibin Cheng, Ming-Ming Shou, Mike Zheng Computer Vision and Pattern Recognition Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approach and present our ImageCritic. We first construct a dataset of reference-degraded-target triplets obtained via VLM-based selection and explicit degradation, which effectively simulates the common inaccuracies or inconsistencies observed in existing generation models. Furthermore, building on a thorough examination of the model's attention mechanisms and intrinsic representations, we accordingly devise an attention alignment loss and a detail encoder to precisely rectify inconsistencies. ImageCritic can be integrated into an agent framework to automatically detect inconsistencies and correct them with multi-round and local editing in complex scenarios. Extensive experiments demonstrate that ImageCritic can effectively resolve detail-related issues in various customized generation scenarios, providing significant improvements over existing methods. |
| title | The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.20614 |