Un-Doubling Diffusion: LLM-guided Disambiguation of Homonym Duplication

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kaskov, Evgeny, Petrova, Elizaveta, Surovtsev, Petr, Kostikova, Anna, Mistiurin, Ilya, Kapitanov, Alexander, Nagaev, Alexander
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915519018303488
author Kaskov, Evgeny
Petrova, Elizaveta
Surovtsev, Petr
Kostikova, Anna
Mistiurin, Ilya
Kapitanov, Alexander
Nagaev, Alexander
author_facet Kaskov, Evgeny
Petrova, Elizaveta
Surovtsev, Petr
Kostikova, Anna
Mistiurin, Ilya
Kapitanov, Alexander
Nagaev, Alexander
contents Homonyms are words with identical spelling but distinct meanings, which pose challenges for many generative models. When a homonym appears in a prompt, diffusion models may generate multiple senses of the word simultaneously, which is known as homonym duplication. This issue is further complicated by an Anglocentric bias, which includes an additional translation step before the text-to-image model pipeline. As a result, even words that are not homonymous in the original language may become homonyms and lose their meaning after translation into English. In this paper, we introduce a method for measuring duplication rates and conduct evaluations of different diffusion models using both automatic evaluation utilizing Vision-Language Models (VLM) and human evaluation. Additionally, we investigate methods to mitigate the homonym duplication problem through prompt expansion, demonstrating that this approach also effectively reduces duplication related to Anglocentric bias. The code for the automatic evaluation pipeline is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Un-Doubling Diffusion: LLM-guided Disambiguation of Homonym Duplication
Kaskov, Evgeny
Petrova, Elizaveta
Surovtsev, Petr
Kostikova, Anna
Mistiurin, Ilya
Kapitanov, Alexander
Nagaev, Alexander
Computation and Language
Homonyms are words with identical spelling but distinct meanings, which pose challenges for many generative models. When a homonym appears in a prompt, diffusion models may generate multiple senses of the word simultaneously, which is known as homonym duplication. This issue is further complicated by an Anglocentric bias, which includes an additional translation step before the text-to-image model pipeline. As a result, even words that are not homonymous in the original language may become homonyms and lose their meaning after translation into English. In this paper, we introduce a method for measuring duplication rates and conduct evaluations of different diffusion models using both automatic evaluation utilizing Vision-Language Models (VLM) and human evaluation. Additionally, we investigate methods to mitigate the homonym duplication problem through prompt expansion, demonstrating that this approach also effectively reduces duplication related to Anglocentric bias. The code for the automatic evaluation pipeline is publicly available.
title Un-Doubling Diffusion: LLM-guided Disambiguation of Homonym Duplication
topic Computation and Language
url https://arxiv.org/abs/2509.21262