Evaluating Perspectival Biases in Cross-Modal Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Saengsukhiran, Teerapol, Chomphooyod, Peerawat, Rodjananant, Narabodee, Chaksangchaichot, Chompakorn, Prakrankamanant, Patawee, Sripheanpol, Witthawin, Lovichit, Pak, Nutanong, Sarana, Chuangsuwanich, Ekapol
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914276368711680
author Saengsukhiran, Teerapol
Chomphooyod, Peerawat
Rodjananant, Narabodee
Chaksangchaichot, Chompakorn
Prakrankamanant, Patawee
Sripheanpol, Witthawin
Lovichit, Pak
Nutanong, Sarana
Chuangsuwanich, Ekapol
author_facet Saengsukhiran, Teerapol
Chomphooyod, Peerawat
Rodjananant, Narabodee
Chaksangchaichot, Chompakorn
Prakrankamanant, Patawee
Sripheanpol, Witthawin
Lovichit, Pak
Nutanong, Sarana
Chuangsuwanich, Ekapol
contents Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: deviations shaped by linguistic prevalence and cultural associations. We introduce the Cross-Cultural, Cross-Modal, Cross-lingual Multimodal (3XCM) benchmark to isolate these effects. Results from our studies indicate that, for image-to-text retrieval, models tend to favor entries from prevalent languages over those that are semantically faithful. For text-to-image retrieval, we observe a consistent "tugging effect" in the joint embedding space between semantic alignment and language-conditioned cultural association. When semantic representations are insufficiently resolved, particularly in low-resource languages, similarity is increasingly governed by culturally familiar visual patterns, leading to systematic association bias in retrieval. Our findings suggest that achieving equitable multimodal retrieval necessitates targeted strategies that explicitly decouple language from culture, rather than relying solely on broader data exposure. This work highlights the need to treat linguistic and cultural biases as distinct, measurable challenges in multimodal representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26861
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Perspectival Biases in Cross-Modal Retrieval
Saengsukhiran, Teerapol
Chomphooyod, Peerawat
Rodjananant, Narabodee
Chaksangchaichot, Chompakorn
Prakrankamanant, Patawee
Sripheanpol, Witthawin
Lovichit, Pak
Nutanong, Sarana
Chuangsuwanich, Ekapol
Information Retrieval
Computation and Language
H.3.3; I.2.7; I.2.10
Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: deviations shaped by linguistic prevalence and cultural associations. We introduce the Cross-Cultural, Cross-Modal, Cross-lingual Multimodal (3XCM) benchmark to isolate these effects. Results from our studies indicate that, for image-to-text retrieval, models tend to favor entries from prevalent languages over those that are semantically faithful. For text-to-image retrieval, we observe a consistent "tugging effect" in the joint embedding space between semantic alignment and language-conditioned cultural association. When semantic representations are insufficiently resolved, particularly in low-resource languages, similarity is increasingly governed by culturally familiar visual patterns, leading to systematic association bias in retrieval. Our findings suggest that achieving equitable multimodal retrieval necessitates targeted strategies that explicitly decouple language from culture, rather than relying solely on broader data exposure. This work highlights the need to treat linguistic and cultural biases as distinct, measurable challenges in multimodal representation learning.
title Evaluating Perspectival Biases in Cross-Modal Retrieval
topic Information Retrieval
Computation and Language
H.3.3; I.2.7; I.2.10
url https://arxiv.org/abs/2510.26861