Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bui, Anh, Vu, Trang, Le, Trung, Kim, Junae, Abraham, Tamas, Omari, Rollin, Kaur, Amar, Phung, Dinh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918354879512576
author Bui, Anh
Vu, Trang
Le, Trung
Kim, Junae
Abraham, Tamas
Omari, Rollin
Kaur, Amar
Phung, Dinh
author_facet Bui, Anh
Vu, Trang
Le, Trung
Kim, Junae
Abraham, Tamas
Omari, Rollin
Kaur, Amar
Phung, Dinh
contents In this paper, we investigate the semantic collapsing problem in generative personalization, an under-explored topic where the learned visual concept ($V$) gradually shifts from its original textual meaning and comes to dominate other concepts in multi-concept input prompts. This issue not only reduces the semantic richness of complex input prompts like "a photo of $V$ wearing glasses and playing guitar" into simpler, less contextually rich forms such as "a photo of $V$" but also leads to simplified output images that fail to capture the intended concept. We identify the root cause as unconstrained optimisation, which allows the learned embedding $V$ to drift arbitrarily in the embedding space, both in direction and magnitude. To address this, we propose a simple yet effective training-free method that adjusts the magnitude and direction of pre-trained embedding at inference time, effectively mitigating the semantic collapsing problem. Our method is broadly applicable across different personalization methods and demonstrates significant improvements in text-image alignment in diverse use cases. Our code is anonymously published at https://github.com/tuananhbui89/Embedding-Adjustment
format Preprint
id arxiv_https___arxiv_org_abs_2506_22685
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
Bui, Anh
Vu, Trang
Le, Trung
Kim, Junae
Abraham, Tamas
Omari, Rollin
Kaur, Amar
Phung, Dinh
Machine Learning
Graphics
In this paper, we investigate the semantic collapsing problem in generative personalization, an under-explored topic where the learned visual concept ($V$) gradually shifts from its original textual meaning and comes to dominate other concepts in multi-concept input prompts. This issue not only reduces the semantic richness of complex input prompts like "a photo of $V$ wearing glasses and playing guitar" into simpler, less contextually rich forms such as "a photo of $V$" but also leads to simplified output images that fail to capture the intended concept. We identify the root cause as unconstrained optimisation, which allows the learned embedding $V$ to drift arbitrarily in the embedding space, both in direction and magnitude. To address this, we propose a simple yet effective training-free method that adjusts the magnitude and direction of pre-trained embedding at inference time, effectively mitigating the semantic collapsing problem. Our method is broadly applicable across different personalization methods and demonstrates significant improvements in text-image alignment in diverse use cases. Our code is anonymously published at https://github.com/tuananhbui89/Embedding-Adjustment
title Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
topic Machine Learning
Graphics
url https://arxiv.org/abs/2506.22685