Saved in:
Bibliographic Details
Main Authors: Li, Yongzhi, Zhang, Saining, Chen, Yibing, Li, Boying, Zhang, Yanxin, Du, Xiaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.07340
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915540828684288
author Li, Yongzhi
Zhang, Saining
Chen, Yibing
Li, Boying
Zhang, Yanxin
Du, Xiaoyu
author_facet Li, Yongzhi
Zhang, Saining
Chen, Yibing
Li, Boying
Zhang, Yanxin
Du, Xiaoyu
contents Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fidelity but are computationally expensive, while learning-based approaches offer efficiency at the cost of entangled representations influenced by nuisance factors. We introduce SpotDiff, a novel learning-based method that extracts subject-specific features by spotting and disentangling interference. Leveraging a pre-trained CLIP image encoder and specialized expert networks for pose and background, SpotDiff isolates subject identity through orthogonality constraints in the feature space. To enable principled training, we introduce SpotDiff10k, a curated dataset with consistent pose and background variations. Experiments demonstrate that SpotDiff achieves more robust subject preservation and controllable editing than prior methods, while attaining competitive performance with only 10k training samples.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
Li, Yongzhi
Zhang, Saining
Chen, Yibing
Li, Boying
Zhang, Yanxin
Du, Xiaoyu
Graphics
Machine Learning
Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fidelity but are computationally expensive, while learning-based approaches offer efficiency at the cost of entangled representations influenced by nuisance factors. We introduce SpotDiff, a novel learning-based method that extracts subject-specific features by spotting and disentangling interference. Leveraging a pre-trained CLIP image encoder and specialized expert networks for pose and background, SpotDiff isolates subject identity through orthogonality constraints in the feature space. To enable principled training, we introduce SpotDiff10k, a curated dataset with consistent pose and background variations. Experiments demonstrate that SpotDiff achieves more robust subject preservation and controllable editing than prior methods, while attaining competitive performance with only 10k training samples.
title SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
topic Graphics
Machine Learning
url https://arxiv.org/abs/2510.07340