An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance
Fuente:
arXiv
Saved in:
| Main Authors: | Khanuja, Simran, Ramamoorthy, Sathyanarayanan, Song, Yueqi, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
Towards Automatic Evaluation for Image Transcreation
by: Khanuja, Simran, et al.
Published: (2024)
by: Khanuja, Simran, et al.
Published: (2024)
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
by: Yayavaram, Arnav, et al.
Published: (2025)
by: Yayavaram, Arnav, et al.
Published: (2025)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
by: Li, Baiqi, et al.
Published: (2024)
by: Li, Baiqi, et al.
Published: (2024)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
by: Ramamoorthy, Sathyanarayanan, et al.
Published: (2025)
by: Ramamoorthy, Sathyanarayanan, et al.
Published: (2025)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
by: Wada, Yuiga, et al.
Published: (2025)
by: Wada, Yuiga, et al.
Published: (2025)
Harnessing Webpage UIs for Text-Rich Visual Understanding
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
Medical thinking with multiple images
by: Yao, Zonghai, et al.
Published: (2026)
by: Yao, Zonghai, et al.
Published: (2026)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
Scaling medical imaging report generation with multimodal reinforcement learning
by: Liu, Qianchu, et al.
Published: (2026)
by: Liu, Qianchu, et al.
Published: (2026)
AutoPresent: Designing Structured Visuals from Scratch
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
A large-scale image-text dataset benchmark for farmland segmentation
by: Tao, Chao, et al.
Published: (2025)
by: Tao, Chao, et al.
Published: (2025)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
by: Onohara, Shota, et al.
Published: (2024)
by: Onohara, Shota, et al.
Published: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
by: Zhang, Sheng, et al.
Published: (2023)
by: Zhang, Sheng, et al.
Published: (2023)
VIMI: Grounding Video Generation through Multi-modal Instruction
by: Fang, Yuwei, et al.
Published: (2024)
by: Fang, Yuwei, et al.
Published: (2024)
Using LLMs as prompt modifier to avoid biases in AI image generators
by: Peinl, René
Published: (2025)
by: Peinl, René
Published: (2025)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
by: Yoon, Yejun, et al.
Published: (2024)
by: Yoon, Yejun, et al.
Published: (2024)
Self-supervised Multi-actor Social Activity Understanding in Streaming Videos
by: Trehan, Shubham, et al.
Published: (2024)
by: Trehan, Shubham, et al.
Published: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
by: Du, Shian, et al.
Published: (2024)
by: Du, Shian, et al.
Published: (2024)
Reverse Stable Diffusion: What prompt was used to generate this image?
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
by: Croitoru, Florinel-Alin, et al.
Published: (2023)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Advancements and limitations of LLMs in replicating human color-word associations
by: Fukushima, Makoto, et al.
Published: (2024)
by: Fukushima, Makoto, et al.
Published: (2024)
MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
by: Li, Jiachun, et al.
Published: (2026)
by: Li, Jiachun, et al.
Published: (2026)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
by: Albadarneh, Israa A., et al.
Published: (2025)
by: Albadarneh, Israa A., et al.
Published: (2025)
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
by: Qi, Xianbiao, et al.
Published: (2025)
by: Qi, Xianbiao, et al.
Published: (2025)
Lacking Data? No worries! How synthetic images can alleviate image scarcity in wildlife surveys: a case study with muskox (Ovibos moschatus)
by: Durand, Simon, et al.
Published: (2025)
by: Durand, Simon, et al.
Published: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
by: Lin, Zhiqiu, et al.
Published: (2024)
by: Lin, Zhiqiu, et al.
Published: (2024)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Large Language Models can Share Images, Too!
by: Lee, Young-Jun, et al.
Published: (2023)
by: Lee, Young-Jun, et al.
Published: (2023)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
by: Tu, Rong-Cheng, et al.
Published: (2024)
by: Tu, Rong-Cheng, et al.
Published: (2024)
Empirical Bayesian image restoration by Langevin sampling with a denoising diffusion implicit prior
by: Mbakam, Charlesquin Kemajou, et al.
Published: (2024)
by: Mbakam, Charlesquin Kemajou, et al.
Published: (2024)
ALGO: Object-Grounded Visual Commonsense Reasoning for Open-World Egocentric Action Recognition
by: Kundu, Sanjoy, et al.
Published: (2024)
by: Kundu, Sanjoy, et al.
Published: (2024)
CRAFT: A Neuro-Symbolic Framework for Visual Functional Affordance Grounding
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
by: Kundu, Sanjoy, et al.
Published: (2025)
by: Kundu, Sanjoy, et al.
Published: (2025)
Hallucinate, Ground, Repeat: A Framework for Generalized Visual Relationship Detection
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
by: Vellamcheti, Shanmukha, et al.
Published: (2025)
Similar Items
-
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
by: Yue, Xiang, et al.
Published: (2024) -
Towards Automatic Evaluation for Image Transcreation
by: Khanuja, Simran, et al.
Published: (2024) -
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
by: Yayavaram, Arnav, et al.
Published: (2025) -
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024) -
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
by: Li, Baiqi, et al.
Published: (2024)