Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mingxiao, Qu, Tingyu, Tuytelaars, Tinne, Moens, Marie-Francine
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910866724618240
author Li, Mingxiao
Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
author_facet Li, Mingxiao
Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
contents Personalized image generation via text prompts has great potential to improve daily life and professional work by facilitating the creation of customized visual content. The aim of image personalization is to create images based on a user-provided subject while maintaining both consistency of the subject and flexibility to accommodate various textual descriptions of that subject. However, current methods face challenges in ensuring fidelity to the text prompt while not overfitting to the training data. In this work, we introduce a novel training pipeline that incorporates an attractor to filter out distractions in training images, allowing the model to focus on learning an effective representation of the personalized subject. Moreover, current evaluation methods struggle due to the lack of a dedicated test set. The evaluation set-up typically relies on the training data of the personalization task to compute text-image and image-image similarity scores, which, while useful, tend to overestimate performance. Although human evaluations are commonly used as an alternative, they often suffer from bias and inconsistency. To address these issues, we curate a diverse and high-quality test set with well-designed prompts. With this new benchmark, automatic evaluation metrics can reliably assess model performance
format Preprint
id arxiv_https___arxiv_org_abs_2503_06632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
Li, Mingxiao
Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
Computer Vision and Pattern Recognition
Personalized image generation via text prompts has great potential to improve daily life and professional work by facilitating the creation of customized visual content. The aim of image personalization is to create images based on a user-provided subject while maintaining both consistency of the subject and flexibility to accommodate various textual descriptions of that subject. However, current methods face challenges in ensuring fidelity to the text prompt while not overfitting to the training data. In this work, we introduce a novel training pipeline that incorporates an attractor to filter out distractions in training images, allowing the model to focus on learning an effective representation of the personalized subject. Moreover, current evaluation methods struggle due to the lack of a dedicated test set. The evaluation set-up typically relies on the training data of the personalization task to compute text-image and image-image similarity scores, which, while useful, tend to overestimate performance. Although human evaluations are commonly used as an alternative, they often suffer from bias and inconsistency. To address these issues, we curate a diverse and high-quality test set with well-designed prompts. With this new benchmark, automatic evaluation metrics can reliably assess model performance
title Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06632