The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ouyang, Ziheng, Song, Yiren, Liu, Yaoli, Zhu, Shihao, Hou, Qibin, Cheng, Ming-Ming, Shou, Mike Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908675966238720
author Ouyang, Ziheng
Song, Yiren
Liu, Yaoli
Zhu, Shihao
Hou, Qibin
Cheng, Ming-Ming
Shou, Mike Zheng
author_facet Ouyang, Ziheng
Song, Yiren
Liu, Yaoli
Zhu, Shihao
Hou, Qibin
Cheng, Ming-Ming
Shou, Mike Zheng
contents Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approach and present our ImageCritic. We first construct a dataset of reference-degraded-target triplets obtained via VLM-based selection and explicit degradation, which effectively simulates the common inaccuracies or inconsistencies observed in existing generation models. Furthermore, building on a thorough examination of the model's attention mechanisms and intrinsic representations, we accordingly devise an attention alignment loss and a detail encoder to precisely rectify inconsistencies. ImageCritic can be integrated into an agent framework to automatically detect inconsistencies and correct them with multi-round and local editing in complex scenarios. Extensive experiments demonstrate that ImageCritic can effectively resolve detail-related issues in various customized generation scenarios, providing significant improvements over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
Ouyang, Ziheng
Song, Yiren
Liu, Yaoli
Zhu, Shihao
Hou, Qibin
Cheng, Ming-Ming
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Previous works have explored various customized generation tasks given a reference image, but they still face limitations in generating consistent fine-grained details. In this paper, our aim is to solve the inconsistency problem of generated images by applying a reference-guided post-editing approach and present our ImageCritic. We first construct a dataset of reference-degraded-target triplets obtained via VLM-based selection and explicit degradation, which effectively simulates the common inaccuracies or inconsistencies observed in existing generation models. Furthermore, building on a thorough examination of the model's attention mechanisms and intrinsic representations, we accordingly devise an attention alignment loss and a detail encoder to precisely rectify inconsistencies. ImageCritic can be integrated into an agent framework to automatically detect inconsistencies and correct them with multi-round and local editing in complex scenarios. Extensive experiments demonstrate that ImageCritic can effectively resolve detail-related issues in various customized generation scenarios, providing significant improvements over existing methods.
title The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20614