Towards Automatic Evaluation for Image Transcreation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khanuja, Simran, Iyer, Vivek, He, Claire, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025)
An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
von: Zhou, Li, et al.
Veröffentlicht: (2025)
von: Zhou, Li, et al.
Veröffentlicht: (2025)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
von: Zhao, Yuming, et al.
Veröffentlicht: (2026)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
von: Wada, Yuiga, et al.
Veröffentlicht: (2025)
von: Wada, Yuiga, et al.
Veröffentlicht: (2025)
Evaluating Text-to-Visual Generation with Image-to-Text Generation
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2024)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Harnessing Webpage UIs for Text-Rich Visual Understanding
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
von: Liu, Junpeng, et al.
Veröffentlicht: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Huixuan, et al.
Veröffentlicht: (2025)
Disability Representations: Finding Biases in Automatic Image Generation
von: Tevissen, Yannis
Veröffentlicht: (2024)
von: Tevissen, Yannis
Veröffentlicht: (2024)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
von: Siingh, Shikhhar, et al.
Veröffentlicht: (2025)
von: Siingh, Shikhhar, et al.
Veröffentlicht: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
von: Matsuda, Kazuki, et al.
Veröffentlicht: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
von: Akhtar, Mubashara, et al.
Veröffentlicht: (2023)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
von: Nyandwi, Jean de Dieu, et al.
Veröffentlicht: (2025)
von: Nyandwi, Jean de Dieu, et al.
Veröffentlicht: (2025)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
von: Liu, Zheng, et al.
Veröffentlicht: (2024)
PRIM: Towards Practical In-Image Multilingual Machine Translation
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
VIMI: Grounding Video Generation through Multi-modal Instruction
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
ELITE: Enhanced Language-Image Toxicity Evaluation for Safety
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
von: Qian, Yusu, et al.
Veröffentlicht: (2024)
Thinking with Programming Vision: Towards a Unified View for Thinking with Images
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
von: Guo, Zirun, et al.
Veröffentlicht: (2025)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
Towards Few-shot Entity Recognition in Document Images: A Graph Neural Network Approach Robust to Image Manipulation
von: Krishnan, Prashant, et al.
Veröffentlicht: (2023)
von: Krishnan, Prashant, et al.
Veröffentlicht: (2023)
A Unified Agentic Framework for Evaluating Conditional Image Generation
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
von: Wang, Jifang, et al.
Veröffentlicht: (2025)
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
von: Geigle, Gregor, et al.
Veröffentlicht: (2023)
Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
von: Yayavaram, Arnav, et al.
Veröffentlicht: (2025) -
An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance
von: Khanuja, Simran, et al.
Veröffentlicht: (2024) -
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024) -
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
von: Yue, Xiang, et al.
Veröffentlicht: (2024) -
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
von: Zhou, Li, et al.
Veröffentlicht: (2025)