Beyond Training for Cultural Awareness: The Role of Dataset Linguistic Structure in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Masoud, Reem I., Feng, Chen, Asano, Shunta, Alshahrani, Saied, Treleaven, Philip Colin, Rodrigues, Miguel R. D. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Cultural Alignment in Large Language Models Using Soft Prompt Tuning
par: Masoud, Reem I., et autres
Publié: (2025)
par: Masoud, Reem I., et autres
Publié: (2025)
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
par: Masoud, Reem I., et autres
Publié: (2023)
par: Masoud, Reem I., et autres
Publié: (2023)
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
par: Alkhowaiter, Mohammed, et autres
Publié: (2025)
par: Alkhowaiter, Mohammed, et autres
Publié: (2025)
Arabic Synonym BERT-based Adversarial Examples for Text Classification
par: Alshahrani, Norah, et autres
Publié: (2024)
par: Alshahrani, Norah, et autres
Publié: (2024)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
par: Asano, Shunta, et autres
Publié: (2026)
par: Asano, Shunta, et autres
Publié: (2026)
CIDAR: Culturally Relevant Instruction Dataset For Arabic
par: Alyafeai, Zaid, et autres
Publié: (2024)
par: Alyafeai, Zaid, et autres
Publié: (2024)
Leveraging Corpus Metadata to Detect Template-based Translation: An Exploratory Case Study of the Egyptian Arabic Wikipedia Edition
par: Alshahrani, Saied, et autres
Publié: (2024)
par: Alshahrani, Saied, et autres
Publié: (2024)
BERT vs GPT for financial engineering
par: Sharkey, Edward, et autres
Publié: (2024)
par: Sharkey, Edward, et autres
Publié: (2024)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
par: Keisha, Figarri, et autres
Publié: (2025)
par: Keisha, Figarri, et autres
Publié: (2025)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
par: Liang, Mengfei, et autres
Publié: (2024)
par: Liang, Mengfei, et autres
Publié: (2024)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
par: Alwajih, Fakhraddin, et autres
Publié: (2025)
par: Alwajih, Fakhraddin, et autres
Publié: (2025)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
par: Handa, Gunmay, et autres
Publié: (2025)
par: Handa, Gunmay, et autres
Publié: (2025)
The Impact of Role Design in In-Context Learning for Large Language Models
par: Rouzegar, Hamidreza, et autres
Publié: (2025)
par: Rouzegar, Hamidreza, et autres
Publié: (2025)
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
par: Kim, Eunsu, et autres
Publié: (2024)
par: Kim, Eunsu, et autres
Publié: (2024)
The Shrinking Landscape of Linguistic Diversity in the Age of Large Language Models
par: Sourati, Zhivar, et autres
Publié: (2025)
par: Sourati, Zhivar, et autres
Publié: (2025)
Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs)
par: Rahman, Abrar, et autres
Publié: (2024)
par: Rahman, Abrar, et autres
Publié: (2024)
Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation
par: Yuan, Zekun, et autres
Publié: (2026)
par: Yuan, Zekun, et autres
Publié: (2026)
Survey of Cultural Awareness in Language Models: Text and Beyond
par: Pawar, Siddhesh, et autres
Publié: (2024)
par: Pawar, Siddhesh, et autres
Publié: (2024)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
par: Kabir, Daeen, et autres
Publié: (2025)
par: Kabir, Daeen, et autres
Publié: (2025)
ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics
par: Al-Khalifa, Hend, et autres
Publié: (2026)
par: Al-Khalifa, Hend, et autres
Publié: (2026)
Decomposed Prompting: Probing Multilingual Linguistic Structure Knowledge in Large Language Models
par: Nie, Ercong, et autres
Publié: (2024)
par: Nie, Ercong, et autres
Publié: (2024)
NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
par: Mekki, Abdellah El, et autres
Publié: (2025)
par: Mekki, Abdellah El, et autres
Publié: (2025)
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
par: King, Theo, et autres
Publié: (2024)
par: King, Theo, et autres
Publié: (2024)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
par: Alwajih, Fakhraddin, et autres
Publié: (2025)
par: Alwajih, Fakhraddin, et autres
Publié: (2025)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
par: Zhou, Xinyu, et autres
Publié: (2024)
par: Zhou, Xinyu, et autres
Publié: (2024)
Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
par: Li, Qiyuan, et autres
Publié: (2025)
par: Li, Qiyuan, et autres
Publié: (2025)
Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language Models
par: Zhang, Binchi, et autres
Publié: (2026)
par: Zhang, Binchi, et autres
Publié: (2026)
Identity-Aware Large Language Models require Cultural Reasoning
par: Plum, Alistair, et autres
Publié: (2025)
par: Plum, Alistair, et autres
Publié: (2025)
Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
par: Xu, Lingnan, et autres
Publié: (2025)
par: Xu, Lingnan, et autres
Publié: (2025)
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models
par: Zhao, Wenlong, et autres
Publié: (2024)
par: Zhao, Wenlong, et autres
Publié: (2024)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
par: Huang, Weiyu, et autres
Publié: (2025)
par: Huang, Weiyu, et autres
Publié: (2025)
On the Compatibility of Generative AI and Generative Linguistics
par: Portelance, Eva, et autres
Publié: (2024)
par: Portelance, Eva, et autres
Publié: (2024)
Linguistic Intelligence in Large Language Models for Telecommunications
par: Ahmed, Tasnim, et autres
Publié: (2024)
par: Ahmed, Tasnim, et autres
Publié: (2024)
Unveiling Linguistic Regions in Large Language Models
par: Zhang, Zhihao, et autres
Publié: (2024)
par: Zhang, Zhihao, et autres
Publié: (2024)
Benchmarking Linguistic Diversity of Large Language Models
par: Guo, Yanzhu, et autres
Publié: (2024)
par: Guo, Yanzhu, et autres
Publié: (2024)
Token-Level Privacy in Large Language Models
par: Harel, Re'em, et autres
Publié: (2025)
par: Harel, Re'em, et autres
Publié: (2025)
Structure-Aware Decoding Mechanisms for Complex Entity Extraction with Large-Scale Language Models
par: Qiu, Zhimin, et autres
Publié: (2025)
par: Qiu, Zhimin, et autres
Publié: (2025)
BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
par: Haque, Farah Binta, et autres
Publié: (2025)
par: Haque, Farah Binta, et autres
Publié: (2025)
DAST: Difficulty-Aware Self-Training on Large Language Models
par: Xue, Boyang, et autres
Publié: (2025)
par: Xue, Boyang, et autres
Publié: (2025)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
par: Wicaksono, Ilham, et autres
Publié: (2025)
par: Wicaksono, Ilham, et autres
Publié: (2025)
Documents similaires
-
Cultural Alignment in Large Language Models Using Soft Prompt Tuning
par: Masoud, Reem I., et autres
Publié: (2025) -
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
par: Masoud, Reem I., et autres
Publié: (2023) -
Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations
par: Alkhowaiter, Mohammed, et autres
Publié: (2025) -
Arabic Synonym BERT-based Adversarial Examples for Text Classification
par: Alshahrani, Norah, et autres
Publié: (2024) -
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
par: Asano, Shunta, et autres
Publié: (2026)