Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Singhania, Aditi, Malani, Krutik, Dhawan, Riddhi, Jain, Arushi, Tandon, Garv, Sharma, Nippun, Chakraborty, Souymodip, Batra, Vineet, Phogat, Ankit
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915611906408448
author Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
author_facet Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
contents Evaluating identity preservation in generative models remains a critical yet unresolved challenge. Existing metrics rely on global embeddings or coarse VLM prompting, failing to capture fine-grained identity changes and providing limited diagnostic insight. We introduce Beyond the Pixels, a hierarchical evaluation framework that decomposes identity assessment into feature-level transformations. Our approach guides VLMs through structured reasoning by (1) hierarchically decomposing subjects into (type, style) -> attribute -> feature decision tree, and (2) prompting for concrete transformations rather than abstract similarity scores. This decomposition grounds VLM analysis in verifiable visual evidence, reducing hallucinations and improving consistency. We validate our framework across four state-of-the-art generative models, demonstrating strong alignment with human judgments in measuring identity consistency. Additionally, we introduce a new benchmark specifically designed to stress-test generative models. It comprises 1,078 image-prompt pairs spanning diverse subject types, including underrepresented categories such as anthropomorphic and animated characters, and captures an average of six to seven transformation axes per prompt.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08087
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
Singhania, Aditi
Malani, Krutik
Dhawan, Riddhi
Jain, Arushi
Tandon, Garv
Sharma, Nippun
Chakraborty, Souymodip
Batra, Vineet
Phogat, Ankit
Computer Vision and Pattern Recognition
Artificial Intelligence
Evaluating identity preservation in generative models remains a critical yet unresolved challenge. Existing metrics rely on global embeddings or coarse VLM prompting, failing to capture fine-grained identity changes and providing limited diagnostic insight. We introduce Beyond the Pixels, a hierarchical evaluation framework that decomposes identity assessment into feature-level transformations. Our approach guides VLMs through structured reasoning by (1) hierarchically decomposing subjects into (type, style) -> attribute -> feature decision tree, and (2) prompting for concrete transformations rather than abstract similarity scores. This decomposition grounds VLM analysis in verifiable visual evidence, reducing hallucinations and improving consistency. We validate our framework across four state-of-the-art generative models, demonstrating strong alignment with human judgments in measuring identity consistency. Additionally, we introduce a new benchmark specifically designed to stress-test generative models. It comprises 1,078 image-prompt pairs spanning diverse subject types, including underrepresented categories such as anthropomorphic and animated characters, and captures an average of six to seven transformation axes per prompt.
title Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.08087