MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bartkowiak, Patryk, Petersen, Lennart, Kotrys, Bartosz, Michels, Dominik, Pirk, Soren, Palubicki, Wojtek
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913152731447296
author Bartkowiak, Patryk
Petersen, Lennart
Kotrys, Bartosz
Michels, Dominik
Pirk, Soren
Palubicki, Wojtek
author_facet Bartkowiak, Patryk
Petersen, Lennart
Kotrys, Bartosz
Michels, Dominik
Pirk, Soren
Palubicki, Wojtek
contents Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a reference, and prompt following, which captures whether the generated scene matches the prompt. Existing metrics commonly compute these signals using global image or text-image embeddings, such as CLIP-I, DINO, and CLIP-T. We show that such metrics correlate poorly with human perception because they attend to the image as a whole instead of separating the concept subject from the background. We introduce MaSC, a masked similarity metric that uses externally provided foreground concept masks to decompose evaluation into subject-specific concept preservation and background-based prompt following. MaSC computes both scores from frozen SigLIP2 SO400M-NaFlex features: concept preservation is measured by masked max-cosine matching between foreground reference patches and generated-image patches, while prompt following is measured by comparing a background-only pooled image embedding to a subject-stripped prompt embedding. On DreamBench++ human ratings, MaSC achieves Krippendorff alpha = 0.471 for concept preservation, outperforming all tested non-LLM baselines and GPT-4V, and approaching GPT-4o. On ORIDa, a real-photo identity-preservation benchmark across physical environments, MaSC achieves AUC = 0.992, nearly perfectly distinguishing same-subject from cross-subject pairs. Its prompt-following score also outperforms the CLIP-T baseline shipped with DreamBench++. These results show that spatially decomposed aggregation is a strong design principle for evaluating concept-driven generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22469
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
Bartkowiak, Patryk
Petersen, Lennart
Kotrys, Bartosz
Michels, Dominik
Pirk, Soren
Palubicki, Wojtek
Computer Vision and Pattern Recognition
I.2.6; I.2.10; I.4.8; I.4.9
Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a reference, and prompt following, which captures whether the generated scene matches the prompt. Existing metrics commonly compute these signals using global image or text-image embeddings, such as CLIP-I, DINO, and CLIP-T. We show that such metrics correlate poorly with human perception because they attend to the image as a whole instead of separating the concept subject from the background. We introduce MaSC, a masked similarity metric that uses externally provided foreground concept masks to decompose evaluation into subject-specific concept preservation and background-based prompt following. MaSC computes both scores from frozen SigLIP2 SO400M-NaFlex features: concept preservation is measured by masked max-cosine matching between foreground reference patches and generated-image patches, while prompt following is measured by comparing a background-only pooled image embedding to a subject-stripped prompt embedding. On DreamBench++ human ratings, MaSC achieves Krippendorff alpha = 0.471 for concept preservation, outperforming all tested non-LLM baselines and GPT-4V, and approaching GPT-4o. On ORIDa, a real-photo identity-preservation benchmark across physical environments, MaSC achieves AUC = 0.992, nearly perfectly distinguishing same-subject from cross-subject pairs. Its prompt-following score also outperforms the CLIP-T baseline shipped with DreamBench++. These results show that spatially decomposed aggregation is a strong design principle for evaluating concept-driven generation.
title MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
topic Computer Vision and Pattern Recognition
I.2.6; I.2.10; I.4.8; I.4.9
url https://arxiv.org/abs/2605.22469