FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Fatehkia, Masoomali, Altinisik, Enes, Sencar, Husrev Taha
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911283529383936
author Fatehkia, Masoomali
Altinisik, Enes
Sencar, Husrev Taha
author_facet Fatehkia, Masoomali
Altinisik, Enes
Sencar, Husrev Taha
contents Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of over 468K prompt and response pairs, drawn from synthetic and public datasets, scored by a panel of LLM judges on harmlessness and cultural awareness, and use it to train two filter variants. To rigorously evaluate cultural alignment, we further develop the first benchmark targeting Arabic cultural contexts, comprising over 1k norm-sensitive prompts with LLM-generated responses annotated by human raters. Results show that FanarGuard achieves stronger agreement with human annotations than inter-annotator reliability, while matching the performance of state-of-the-art filters on safety benchmarks. These findings highlight the importance of integrating cultural awareness into moderation and establish FanarGuard as a practical step toward more context-sensitive safeguards.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
Fatehkia, Masoomali
Altinisik, Enes
Sencar, Husrev Taha
Computation and Language
Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on general safety and overlook cultural context. In this work, we introduce FanarGuard, a bilingual moderation filter that evaluates both safety and cultural alignment in Arabic and English. We construct a dataset of over 468K prompt and response pairs, drawn from synthetic and public datasets, scored by a panel of LLM judges on harmlessness and cultural awareness, and use it to train two filter variants. To rigorously evaluate cultural alignment, we further develop the first benchmark targeting Arabic cultural contexts, comprising over 1k norm-sensitive prompts with LLM-generated responses annotated by human raters. Results show that FanarGuard achieves stronger agreement with human annotations than inter-annotator reliability, while matching the performance of state-of-the-art filters on safety benchmarks. These findings highlight the importance of integrating cultural awareness into moderation and establish FanarGuard as a practical step toward more context-sensitive safeguards.
title FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
topic Computation and Language
url https://arxiv.org/abs/2511.18852