Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kabir, Mohsinul, Abrar, Ajwad, Ananiadou, Sophia
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908649239085056
author Kabir, Mohsinul
Abrar, Ajwad
Ananiadou, Sophia
author_facet Kabir, Mohsinul
Abrar, Ajwad
Ananiadou, Sophia
contents A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more realistic, unconstrained approaches. Using the World Values Survey (WVS) and Hofstede Cultural Dimensions as case studies, we demonstrate that LLMs exhibit stronger cultural alignment in less constrained settings, where responses are not forced. Additionally, we show that even minor changes, such as reordering survey choices, lead to inconsistent outputs, exposing the limitations of closed-style evaluations. Our findings advocate for more robust and flexible evaluation frameworks that focus on specific cultural proxies, encouraging more nuanced and accurate assessments of cultural alignment in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
Kabir, Mohsinul
Abrar, Ajwad
Ananiadou, Sophia
Computation and Language
Artificial Intelligence
Computers and Society
A large number of studies rely on closed-style multiple-choice surveys to evaluate cultural alignment in Large Language Models (LLMs). In this work, we challenge this constrained evaluation paradigm and explore more realistic, unconstrained approaches. Using the World Values Survey (WVS) and Hofstede Cultural Dimensions as case studies, we demonstrate that LLMs exhibit stronger cultural alignment in less constrained settings, where responses are not forced. Additionally, we show that even minor changes, such as reordering survey choices, lead to inconsistent outputs, exposing the limitations of closed-style evaluations. Our findings advocate for more robust and flexible evaluation frameworks that focus on specific cultural proxies, encouraging more nuanced and accurate assessments of cultural alignment in LLMs.
title Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2502.08045