Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Chuancheng, Li, Shangze, Guo, Shiming, Xie, Simiao, Wu, Wenhua, Dou, Jingtong, Wu, Chao, Xiao, Canran, Wang, Cong, Cheng, Zifeng, Shen, Fei, Chua, Tat-Seng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915630999928832
author Shi, Chuancheng
Li, Shangze
Guo, Shiming
Xie, Simiao
Wu, Wenhua
Dou, Jingtong
Wu, Chao
Xiao, Canran
Wang, Cong
Cheng, Zifeng
Shen, Fei
Chua, Tat-Seng
author_facet Shi, Chuancheng
Li, Shangze
Guo, Shiming
Xie, Simiao
Wu, Wenhua
Dou, Jingtong
Wu, Chao
Xiao, Canran
Wang, Cong
Cheng, Zifeng
Shen, Fei
Chua, Tat-Seng
contents Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-lingual cultural consistency. We conduct a comprehensive analysis showing that current T2I models often produce culturally neutral or English-biased results under multilingual prompts. Analyses of two representative models indicate that the issue stems not from missing cultural knowledge but from insufficient activation of culture-related representations. We propose a probing method that localizes culture-sensitive signals to a small set of neurons in a few fixed layers. Guided by this finding, we introduce two complementary alignment strategies: (1) inference-time cultural activation that amplifies the identified neurons without backbone fine-tuned; and (2) layer-targeted cultural enhancement that updates only culturally relevant layers. Experiments on our CultureBench demonstrate consistent improvements over strong baselines in cultural consistency while preserving fidelity and diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2511_17282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
Shi, Chuancheng
Li, Shangze
Guo, Shiming
Xie, Simiao
Wu, Wenhua
Dou, Jingtong
Wu, Chao
Xiao, Canran
Wang, Cong
Cheng, Zifeng
Shen, Fei
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-lingual cultural consistency. We conduct a comprehensive analysis showing that current T2I models often produce culturally neutral or English-biased results under multilingual prompts. Analyses of two representative models indicate that the issue stems not from missing cultural knowledge but from insufficient activation of culture-related representations. We propose a probing method that localizes culture-sensitive signals to a small set of neurons in a few fixed layers. Guided by this finding, we introduce two complementary alignment strategies: (1) inference-time cultural activation that amplifies the identified neurons without backbone fine-tuned; and (2) layer-targeted cultural enhancement that updates only culturally relevant layers. Experiments on our CultureBench demonstrate consistent improvements over strong baselines in cultural consistency while preserving fidelity and diversity.
title Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2511.17282