Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
Fuente:
arXiv
Saved in:
| Main Authors: | Gomes, Gonçalo, Zerva, Chrysoula, Martins, Bruno |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
by: Ulmer, Dennis, et al.
Published: (2024)
by: Ulmer, Dennis, et al.
Published: (2024)
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings
by: Gomes, Gonçalo, et al.
Published: (2024)
by: Gomes, Gonçalo, et al.
Published: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
Rejected Dialects: Biases Against African American Language in Reward Models
by: Mire, Joel, et al.
Published: (2025)
by: Mire, Joel, et al.
Published: (2025)
Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
by: Mohammed, Wafaa, et al.
Published: (2025)
by: Mohammed, Wafaa, et al.
Published: (2025)
How and where does CLIP process negation?
by: Quantmeyer, Vincent, et al.
Published: (2024)
by: Quantmeyer, Vincent, et al.
Published: (2024)
Subjective Logic Encodings
by: Vasilakes, Jake, et al.
Published: (2025)
by: Vasilakes, Jake, et al.
Published: (2025)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
by: Yin, Ho, et al.
Published: (2025)
by: Yin, Ho, et al.
Published: (2025)
How to Understand Named Entities: Using Common Sense for News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
ChipGPT: How far are we from natural language hardware design
by: Chang, Kaiyan, et al.
Published: (2023)
by: Chang, Kaiyan, et al.
Published: (2023)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
Counterfactual Fairness with Graph Uncertainty
by: Valério, Davi, et al.
Published: (2026)
by: Valério, Davi, et al.
Published: (2026)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
by: Gupta, Ashray, et al.
Published: (2025)
by: Gupta, Ashray, et al.
Published: (2025)
Can we Evaluate RAGs with Synthetic Data?
by: van Elburg, Jonas, et al.
Published: (2025)
by: van Elburg, Jonas, et al.
Published: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
by: Matsuda, Kazuki, et al.
Published: (2024)
by: Matsuda, Kazuki, et al.
Published: (2024)
How Vocabulary Sharing Facilitates Multilingualism in LLaMA?
by: Yuan, Fei, et al.
Published: (2023)
by: Yuan, Fei, et al.
Published: (2023)
How do Large Language Models Handle Multilingualism?
by: Zhao, Yiran, et al.
Published: (2024)
by: Zhao, Yiran, et al.
Published: (2024)
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants
by: Borges, Beatriz, et al.
Published: (2024)
by: Borges, Beatriz, et al.
Published: (2024)
MELA: Multilingual Evaluation of Linguistic Acceptability
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
by: Oduwole, Mardiyyah, et al.
Published: (2025)
by: Oduwole, Mardiyyah, et al.
Published: (2025)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
by: de Carvalho, Gonçalo Hora, et al.
Published: (2024)
by: de Carvalho, Gonçalo Hora, et al.
Published: (2024)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
The Instruction Gap: LLMs get lost in Following Instruction
by: Tripathi, Vishesh, et al.
Published: (2025)
by: Tripathi, Vishesh, et al.
Published: (2025)
How well can LLMs Grade Essays in Arabic?
by: Ghazawi, Rayed, et al.
Published: (2025)
by: Ghazawi, Rayed, et al.
Published: (2025)
UoR-NCL at SemEval-2025 Task 1: Using Generative LLMs and CLIP Models for Multilingual Multimodal Idiomaticity Representation
by: Markchom, Thanet, et al.
Published: (2025)
by: Markchom, Thanet, et al.
Published: (2025)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
by: Matsuda, Kazuki, et al.
Published: (2025)
by: Matsuda, Kazuki, et al.
Published: (2025)
How does a Multilingual LM Handle Multiple Languages?
by: Kakarla, Santhosh, et al.
Published: (2025)
by: Kakarla, Santhosh, et al.
Published: (2025)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
by: Liu, Danni, et al.
Published: (2023)
by: Liu, Danni, et al.
Published: (2023)
M-IFEval: Multilingual Instruction-Following Evaluation
by: Dussolle, Antoine, et al.
Published: (2025)
by: Dussolle, Antoine, et al.
Published: (2025)
The Roles of English in Evaluating Multilingual Language Models
by: Poelman, Wessel, et al.
Published: (2024)
by: Poelman, Wessel, et al.
Published: (2024)
Translation as a Scalable Proxy for Multilingual Evaluation
by: Issaka, Sheriff, et al.
Published: (2026)
by: Issaka, Sheriff, et al.
Published: (2026)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
by: Chen, Pinzhen, et al.
Published: (2024)
by: Chen, Pinzhen, et al.
Published: (2024)
Unsupervised Flow Discovery from Task-oriented Dialogues
by: Ferreira, Patrícia, et al.
Published: (2024)
by: Ferreira, Patrícia, et al.
Published: (2024)
Similar Items
-
A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
by: Gomes, Gonçalo, et al.
Published: (2025) -
Non-Exchangeable Conformal Language Generation with Nearest Neighbors
by: Ulmer, Dennis, et al.
Published: (2024) -
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings
by: Gomes, Gonçalo, et al.
Published: (2024) -
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024) -
A Survey of LLM-based Agents in Medicine: How far are we from Baymax?
by: Wang, Wenxuan, et al.
Published: (2025)