UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cheng, Yufeng, Wu, Wenxu, Wu, Shaojin, Huang, Mengqi, Ding, Fei, He, Qian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915484345040896
author Cheng, Yufeng
Wu, Wenxu
Wu, Shaojin
Huang, Mengqi
Ding, Fei
He, Qian
author_facet Cheng, Yufeng
Wu, Wenxu
Wu, Shaojin
Huang, Mengqi
Ding, Fei
He, Qian
contents Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-reference images, limiting the identity scalability of customization models. To address this, we present UMO, a Unified Multi-identity Optimization framework, designed to maintain high-fidelity identity preservation and alleviate identity confusion with scalability. With "multi-to-multi matching" paradigm, UMO reformulates multi-identity generation as a global assignment optimization problem and unleashes multi-identity consistency for existing image customization methods generally through reinforcement learning on diffusion models. To facilitate the training of UMO, we develop a scalable customization dataset with multi-reference images, consisting of both synthesised and real parts. Additionally, we propose a new metric to measure identity confusion. Extensive experiments demonstrate that UMO not only improves identity consistency significantly, but also reduces identity confusion on several image customization methods, setting a new state-of-the-art among open-source methods along the dimension of identity preserving. Code and model: https://github.com/bytedance/UMO
format Preprint
id arxiv_https___arxiv_org_abs_2509_06818
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
Cheng, Yufeng
Wu, Wenxu
Wu, Shaojin
Huang, Mengqi
Ding, Fei
He, Qian
Computer Vision and Pattern Recognition
Machine Learning
Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to faces, a significant challenge remains in preserving consistent identity while avoiding identity confusion with multi-reference images, limiting the identity scalability of customization models. To address this, we present UMO, a Unified Multi-identity Optimization framework, designed to maintain high-fidelity identity preservation and alleviate identity confusion with scalability. With "multi-to-multi matching" paradigm, UMO reformulates multi-identity generation as a global assignment optimization problem and unleashes multi-identity consistency for existing image customization methods generally through reinforcement learning on diffusion models. To facilitate the training of UMO, we develop a scalable customization dataset with multi-reference images, consisting of both synthesised and real parts. Additionally, we propose a new metric to measure identity confusion. Extensive experiments demonstrate that UMO not only improves identity consistency significantly, but also reduces identity confusion on several image customization methods, setting a new state-of-the-art among open-source methods along the dimension of identity preserving. Code and model: https://github.com/bytedance/UMO
title UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.06818