Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Masoud, Reem I., Feng, Chen, Asano, Shunta, Alshahrani, Saied, Treleaven, Philip Colin, Rodrigues, Miguel R. D.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911414027812864
author Masoud, Reem I.
Feng, Chen
Asano, Shunta
Alshahrani, Saied
Treleaven, Philip Colin
Rodrigues, Miguel R. D.
author_facet Masoud, Reem I.
Feng, Chen
Asano, Shunta
Alshahrani, Saied
Treleaven, Philip Colin
Rodrigues, Miguel R. D.
contents The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adaptation remain poorly understood. We adopt a dataset-centric view of cultural alignment and ask which linguistic properties of fine-tuning data are associated with cultural performance, whether these properties are predictive prior to training, and how these effects vary across models. We compute lightweight linguistic, semantic, and structural metrics for Arabic, Chinese, and Japanese datasets and apply principal component analysis separately within each language. This design ensures that the resulting components capture variation among datasets written in the same language rather than differences between languages. The resulting components correspond to broadly interpretable axes related to semantic coherence, surface-level lexical and syntactic diversity, and lexical or structural richness, though their composition varies across languages. We fine-tune three major LLM families (LLaMA, Mistral, DeepSeek) and evaluate them on benchmarks of cultural knowledge, values, and norms. While PCA components correlate with downstream performance, these associations are strongly model-dependent. Through controlled subset interventions, we show that lexical-oriented components (PC3) are the most robust, yielding more consistent performance across models and benchmarks, whereas emphasizing semantic or diversity extremes (PC1-PC2) is often neutral or harmful.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01161
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models
Masoud, Reem I.
Feng, Chen
Asano, Shunta
Alshahrani, Saied
Treleaven, Philip Colin
Rodrigues, Miguel R. D.
Computation and Language
The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning datasets used for cultural adaptation remain poorly understood. We adopt a dataset-centric view of cultural alignment and ask which linguistic properties of fine-tuning data are associated with cultural performance, whether these properties are predictive prior to training, and how these effects vary across models. We compute lightweight linguistic, semantic, and structural metrics for Arabic, Chinese, and Japanese datasets and apply principal component analysis separately within each language. This design ensures that the resulting components capture variation among datasets written in the same language rather than differences between languages. The resulting components correspond to broadly interpretable axes related to semantic coherence, surface-level lexical and syntactic diversity, and lexical or structural richness, though their composition varies across languages. We fine-tune three major LLM families (LLaMA, Mistral, DeepSeek) and evaluate them on benchmarks of cultural knowledge, values, and norms. While PCA components correlate with downstream performance, these associations are strongly model-dependent. Through controlled subset interventions, we show that lexical-oriented components (PC3) are the most robust, yielding more consistent performance across models and benchmarks, whereas emphasizing semantic or diversity extremes (PC1-PC2) is often neutral or harmful.
title Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2602.01161