Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singhania, Aditi, Malani, Krutik, Dhawan, Riddhi, Jain, Arushi, Tandon, Garv, Sharma, Nippun, Chakraborty, Souymodip, Batra, Vineet, Phogat, Ankit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915611906408448
author Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
author_facet Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
contents Evaluating identity preservation in generative models remains a critical yet unresolved challenge. Existing metrics rely on global embeddings or coarse VLM prompting, failing to capture fine-grained identity changes and providing limited diagnostic insight. We introduce Beyond the Pixels, a hierarchical evaluation framework that decomposes identity assessment into feature-level transformations. Our approach guides VLMs through structured reasoning by (1) hierarchically decomposing subjects into (type, style) -> attribute -> feature decision tree, and (2) prompting for concrete transformations rather than abstract similarity scores. This decomposition grounds VLM analysis in verifiable visual evidence, reducing hallucinations and improving consistency. We validate our framework across four state-of-the-art generative models, demonstrating strong alignment with human judgments in measuring identity consistency. Additionally, we introduce a new benchmark specifically designed to stress-test generative models. It comprises 1,078 image-prompt pairs spanning diverse subject types, including underrepresented categories such as anthropomorphic and animated characters, and captures an average of six to seven transformation axes per prompt.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
Computer Vision and Pattern Recognition
Artificial Intelligence
Evaluating identity preservation in generative models remains a critical yet unresolved challenge. Existing metrics rely on global embeddings or coarse VLM prompting, failing to capture fine-grained identity changes and providing limited diagnostic insight. We introduce Beyond the Pixels, a hierarchical evaluation framework that decomposes identity assessment into feature-level transformations. Our approach guides VLMs through structured reasoning by (1) hierarchically decomposing subjects into (type, style) -> attribute -> feature decision tree, and (2) prompting for concrete transformations rather than abstract similarity scores. This decomposition grounds VLM analysis in verifiable visual evidence, reducing hallucinations and improving consistency. We validate our framework across four state-of-the-art generative models, demonstrating strong alignment with human judgments in measuring identity consistency. Additionally, we introduce a new benchmark specifically designed to stress-test generative models. It comprises 1,078 image-prompt pairs spanning diverse subject types, including underrepresented categories such as anthropomorphic and animated characters, and captures an average of six to seven transformation axes per prompt.
title Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.08087