Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915630999928832 |
|---|---|
| author | Shi, Chuancheng Li, Shangze Guo, Shiming Xie, Simiao Wu, Wenhua Dou, Jingtong Wu, Chao Xiao, Canran Wang, Cong Cheng, Zifeng Shen, Fei Chua, Tat-Seng |
| author_facet | Shi, Chuancheng Li, Shangze Guo, Shiming Xie, Simiao Wu, Wenhua Dou, Jingtong Wu, Chao Xiao, Canran Wang, Cong Cheng, Zifeng Shen, Fei Chua, Tat-Seng |
| contents | Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-lingual cultural consistency. We conduct a comprehensive analysis showing that current T2I models often produce culturally neutral or English-biased results under multilingual prompts. Analyses of two representative models indicate that the issue stems not from missing cultural knowledge but from insufficient activation of culture-related representations. We propose a probing method that localizes culture-sensitive signals to a small set of neurons in a few fixed layers. Guided by this finding, we introduce two complementary alignment strategies: (1) inference-time cultural activation that amplifies the identified neurons without backbone fine-tuned; and (2) layer-targeted cultural enhancement that updates only culturally relevant layers. Experiments on our CultureBench demonstrate consistent improvements over strong baselines in cultural consistency while preserving fidelity and diversity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_17282 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation Shi, Chuancheng Li, Shangze Guo, Shiming Xie, Simiao Wu, Wenhua Dou, Jingtong Wu, Chao Xiao, Canran Wang, Cong Cheng, Zifeng Shen, Fei Chua, Tat-Seng Computer Vision and Pattern Recognition Artificial Intelligence Computers and Society Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural contexts: because language carries cultural connotations, images synthesized from multilingual prompts should preserve cross-lingual cultural consistency. We conduct a comprehensive analysis showing that current T2I models often produce culturally neutral or English-biased results under multilingual prompts. Analyses of two representative models indicate that the issue stems not from missing cultural knowledge but from insufficient activation of culture-related representations. We propose a probing method that localizes culture-sensitive signals to a small set of neurons in a few fixed layers. Guided by this finding, we introduce two complementary alignment strategies: (1) inference-time cultural activation that amplifies the identified neurons without backbone fine-tuned; and (2) layer-targeted cultural enhancement that updates only culturally relevant layers. Experiments on our CultureBench demonstrate consistent improvements over strong baselines in cultural consistency while preserving fidelity and diversity. |
| title | Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2511.17282 |