When Cultures Meet: Multicultural Text-to-Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhalerao, Parth, Yalamarty, Mounika, Trinh, Brian, Ignat, Oana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
por: Li, Shuowei, et al.
Publicado: (2026)
por: Li, Shuowei, et al.
Publicado: (2026)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
por: Bai, Longju, et al.
Publicado: (2024)
por: Bai, Longju, et al.
Publicado: (2024)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
por: Zhao, Yuming, et al.
Publicado: (2026)
por: Zhao, Yuming, et al.
Publicado: (2026)
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
por: Ignat, Oana, et al.
Publicado: (2024)
por: Ignat, Oana, et al.
Publicado: (2024)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
por: Nwatu, Joan, et al.
Publicado: (2025)
por: Nwatu, Joan, et al.
Publicado: (2025)
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
por: Bhalerao, Parth, et al.
Publicado: (2026)
por: Bhalerao, Parth, et al.
Publicado: (2026)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
por: Jeong, Suchae, et al.
Publicado: (2025)
por: Jeong, Suchae, et al.
Publicado: (2025)
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
por: Adamkiewicz, Krzysztof, et al.
Publicado: (2026)
por: Adamkiewicz, Krzysztof, et al.
Publicado: (2026)
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
por: Shi, Chuancheng, et al.
Publicado: (2025)
por: Shi, Chuancheng, et al.
Publicado: (2025)
When Color-Space Decoupling Meets Diffusion for Adverse-Weather Image Restoration
por: Fang, Wenxuan, et al.
Publicado: (2025)
por: Fang, Wenxuan, et al.
Publicado: (2025)
MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer
por: Ignat, Polezhaev, et al.
Publicado: (2024)
por: Ignat, Polezhaev, et al.
Publicado: (2024)
Enlighten-Your-Voice: When Multimodal Meets Zero-shot Low-light Image Enhancement
por: Zhang, Xiaofeng, et al.
Publicado: (2023)
por: Zhang, Xiaofeng, et al.
Publicado: (2023)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
por: Jang, Young Kyun, et al.
Publicado: (2024)
por: Jang, Young Kyun, et al.
Publicado: (2024)
RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
por: Vasilev, Viacheslav, et al.
Publicado: (2025)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
por: Luo, Junwei, et al.
Publicado: (2025)
por: Luo, Junwei, et al.
Publicado: (2025)
Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
When Text and Images Don't Mix: Bias-Correcting Language-Image Similarity Scores for Anomaly Detection
por: Goodge, Adam, et al.
Publicado: (2024)
por: Goodge, Adam, et al.
Publicado: (2024)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
por: Nwatu, Joan, et al.
Publicado: (2024)
por: Nwatu, Joan, et al.
Publicado: (2024)
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
por: Xu, Haidong, et al.
Publicado: (2024)
por: Xu, Haidong, et al.
Publicado: (2024)
Agentic Retoucher for Text-To-Image Generation
por: Shen, Shaocheng, et al.
Publicado: (2026)
por: Shen, Shaocheng, et al.
Publicado: (2026)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
por: Kim, Junho, et al.
Publicado: (2026)
por: Kim, Junho, et al.
Publicado: (2026)
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
por: Choi, Dasol, et al.
Publicado: (2025)
por: Choi, Dasol, et al.
Publicado: (2025)
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
por: Pang, Zirui, et al.
Publicado: (2025)
por: Pang, Zirui, et al.
Publicado: (2025)
Addressing Image Authenticity When Cameras Use Generative AI
por: Masud, Umar, et al.
Publicado: (2026)
por: Masud, Umar, et al.
Publicado: (2026)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
por: Nayak, Shravan, et al.
Publicado: (2025)
por: Nayak, Shravan, et al.
Publicado: (2025)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
por: Lim, Youngsun, et al.
Publicado: (2024)
por: Lim, Youngsun, et al.
Publicado: (2024)
Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
por: Zhang, Guoqing, et al.
Publicado: (2025)
por: Zhang, Guoqing, et al.
Publicado: (2025)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
por: Poletti, Silvia, et al.
Publicado: (2026)
por: Poletti, Silvia, et al.
Publicado: (2026)
Personalized Reward Modeling for Text-to-Image Generation
por: Lee, Jeongeun, et al.
Publicado: (2025)
por: Lee, Jeongeun, et al.
Publicado: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
por: Mo, Wenyi, et al.
Publicado: (2024)
por: Mo, Wenyi, et al.
Publicado: (2024)
CLAReSNet: When Convolution Meets Latent Attention for Hyperspectral Image Classification
por: Bandyopadhyay, Asmit, et al.
Publicado: (2025)
por: Bandyopadhyay, Asmit, et al.
Publicado: (2025)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
por: Lim, Youngsun, et al.
Publicado: (2024)
por: Lim, Youngsun, et al.
Publicado: (2024)
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
por: Padalkar, Parth, et al.
Publicado: (2025)
por: Padalkar, Parth, et al.
Publicado: (2025)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
por: Sampaio, Georgia Gabriela, et al.
Publicado: (2024)
por: Sampaio, Georgia Gabriela, et al.
Publicado: (2024)
MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data
por: Berman, William, et al.
Publicado: (2024)
por: Berman, William, et al.
Publicado: (2024)
Regeneration Based Training-free Attribution of Fake Images Generated by Text-to-Image Generative Models
por: Li, Meiling, et al.
Publicado: (2024)
por: Li, Meiling, et al.
Publicado: (2024)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
por: Lakhanpal, Sanyam, et al.
Publicado: (2024)
HARIVO: Harnessing Text-to-Image Models for Video Generation
por: Kwon, Mingi, et al.
Publicado: (2024)
por: Kwon, Mingi, et al.
Publicado: (2024)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
por: Vice, Jordan, et al.
Publicado: (2024)
por: Vice, Jordan, et al.
Publicado: (2024)
Interactive Visual Assessment for Text-to-Image Generation Models
por: Mi, Xiaoyue, et al.
Publicado: (2024)
por: Mi, Xiaoyue, et al.
Publicado: (2024)
Ejemplares similares
-
MAVEN A Multi-Agent Framework for Multicultural Text-to-Video Generation
por: Li, Shuowei, et al.
Publicado: (2026) -
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
por: Bai, Longju, et al.
Publicado: (2024) -
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
por: Zhao, Yuming, et al.
Publicado: (2026) -
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost
por: Ignat, Oana, et al.
Publicado: (2024) -
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
por: Nwatu, Joan, et al.
Publicado: (2025)