Probing Cultural Signals in Large Language Models through Author Profiling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lafargue, Valentin, Guerra-Adames, Ariel, Claeys, Emmanuelle, Vuichard, Elouan, Loubes, Jean-Michel
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917353155985408
author Lafargue, Valentin
Guerra-Adames, Ariel
Claeys, Emmanuelle
Vuichard, Elouan
Loubes, Jean-Michel
author_facet Lafargue, Valentin
Guerra-Adames, Ariel
Claeys, Emmanuelle
Vuichard, Elouan
Loubes, Jean-Michel
contents Large language models (LLMs) are increasingly deployed in applications with societal impact, raising concerns about the cultural biases they encode. We probe these representations by evaluating whether LLMs can perform author profiling from song lyrics in a zero-shot setting, inferring singers' gender and ethnicity without task-specific fine-tuning. Across several open-source models evaluated on more than 10,000 lyrics, we find that LLMs achieve non-trivial profiling performance but demonstrate systematic cultural alignment: most models default toward North American ethnicity, while DeepSeek-1.5B aligns more strongly with Asian ethnicity. This finding emerges from both the models' prediction distributions and an analysis of their generated rationales. To quantify these disparities, we introduce two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), and show that Ministral-8B displays the strongest ethnicity bias among the evaluated models, whereas Gemma-12B shows the most balanced behavior. Our code is available on [GitHub](https://github.com/ValentinLafargue/CulturalProbingLLM) and results on [HuggingFace](https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults).
format Preprint
id arxiv_https___arxiv_org_abs_2603_16749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing Cultural Signals in Large Language Models through Author Profiling
Lafargue, Valentin
Guerra-Adames, Ariel
Claeys, Emmanuelle
Vuichard, Elouan
Loubes, Jean-Michel
Computation and Language
Machine Learning
Large language models (LLMs) are increasingly deployed in applications with societal impact, raising concerns about the cultural biases they encode. We probe these representations by evaluating whether LLMs can perform author profiling from song lyrics in a zero-shot setting, inferring singers' gender and ethnicity without task-specific fine-tuning. Across several open-source models evaluated on more than 10,000 lyrics, we find that LLMs achieve non-trivial profiling performance but demonstrate systematic cultural alignment: most models default toward North American ethnicity, while DeepSeek-1.5B aligns more strongly with Asian ethnicity. This finding emerges from both the models' prediction distributions and an analysis of their generated rationales. To quantify these disparities, we introduce two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), and show that Ministral-8B displays the strongest ethnicity bias among the evaluated models, whereas Gemma-12B shows the most balanced behavior. Our code is available on [GitHub](https://github.com/ValentinLafargue/CulturalProbingLLM) and results on [HuggingFace](https://huggingface.co/datasets/ValentinLAFARGUE/AuthorProfilingResults).
title Probing Cultural Signals in Large Language Models through Author Profiling
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.16749