No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mohamed, Youssef, Li, Runjia, Ahmad, Ibrahim Said, Haydarov, Kilichbek, Torr, Philip, Church, Kenneth Ward, Elhoseiny, Mohamed
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912107131305984
author Mohamed, Youssef
Li, Runjia
Ahmad, Ibrahim Said
Haydarov, Kilichbek
Torr, Philip
Church, Kenneth Ward
Elhoseiny, Mohamed
author_facet Mohamed, Youssef
Li, Runjia
Ahmad, Ibrahim Said
Haydarov, Kilichbek
Torr, Philip
Church, Kenneth Ward
Elhoseiny, Mohamed
contents Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjective emotions and ArtELingo introduced some multilinguality (Chinese and Arabic). However we believe there should be more multilinguality. Hence, we present ArtELingo-28, a vision-language benchmark that spans $\textbf{28}$ languages and encompasses approximately $\textbf{200,000}$ annotations ($\textbf{140}$ annotations per image). Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo-28 emphasizes diversity of opinions over languages and cultures. The challenge is to build machine learning systems that assign emotional captions to images. Baseline results will be presented for three novel conditions: Zero-Shot, Few-Shot and One-vs-All Zero-Shot. We find that cross-lingual transfer is more successful for culturally-related languages. Data and code are provided at www.artelingo.org.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03769
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
Mohamed, Youssef
Li, Runjia
Ahmad, Ibrahim Said
Haydarov, Kilichbek
Torr, Philip
Church, Kenneth Ward
Elhoseiny, Mohamed
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
Research in vision and language has made considerable progress thanks to benchmarks such as COCO. COCO captions focused on unambiguous facts in English; ArtEmis introduced subjective emotions and ArtELingo introduced some multilinguality (Chinese and Arabic). However we believe there should be more multilinguality. Hence, we present ArtELingo-28, a vision-language benchmark that spans $\textbf{28}$ languages and encompasses approximately $\textbf{200,000}$ annotations ($\textbf{140}$ annotations per image). Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo-28 emphasizes diversity of opinions over languages and cultures. The challenge is to build machine learning systems that assign emotional captions to images. Baseline results will be presented for three novel conditions: Zero-Shot, Few-Shot and One-vs-All Zero-Shot. We find that cross-lingual transfer is more successful for culturally-related languages. Data and code are provided at www.artelingo.org.
title No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2411.03769