FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Fatehkia, Masoomali, Altinisik, Enes, Sencar, Husrev Taha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAM: Training Policy-Aligned Moderation Filters at Scale
by: Fatehkia, Masoomali, et al.
Published: (2025)
by: Fatehkia, Masoomali, et al.
Published: (2025)
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
by: Ersoy, Asim, et al.
Published: (2025)
by: Ersoy, Asim, et al.
Published: (2025)
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
by: Altinisik, Enes, et al.
Published: (2026)
by: Altinisik, Enes, et al.
Published: (2026)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024)
by: Altinisik, Enes, et al.
Published: (2024)
Fanar 2.0: Arabic Generative AI Stack
by: FANAR TEAM, et al.
Published: (2026)
by: FANAR TEAM, et al.
Published: (2026)
Fanar: An Arabic-Centric Multimodal Generative AI Platform
by: Fanar Team, et al.
Published: (2025)
by: Fanar Team, et al.
Published: (2025)
There Is More to Refusal in Large Language Models than a Single Direction
by: Joad, Faaiz, et al.
Published: (2026)
by: Joad, Faaiz, et al.
Published: (2026)
T-RAG: Lessons from the LLM Trenches
by: Fatehkia, Masoomali, et al.
Published: (2024)
by: Fatehkia, Masoomali, et al.
Published: (2024)
Multimedia Forensics
by: Husrev Taha Sencar
by: Husrev Taha Sencar
CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
by: Joshi, Raviraj, et al.
Published: (2025)
by: Joshi, Raviraj, et al.
Published: (2025)
From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA
by: Abbas, Ummar, et al.
Published: (2026)
by: Abbas, Ummar, et al.
Published: (2026)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
by: Kumar, Priyanshu, et al.
Published: (2025)
by: Kumar, Priyanshu, et al.
Published: (2025)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
STAND-Guard: A Small Task-Adaptive Content Moderation Model
by: Wang, Minjia, et al.
Published: (2024)
by: Wang, Minjia, et al.
Published: (2024)
ExpGuard: LLM Content Moderation in Specialized Domains
by: Choi, Minseok, et al.
Published: (2026)
by: Choi, Minseok, et al.
Published: (2026)
Command R7B Arabic: A Small, Enterprise Focused, Multilingual, and Culturally Aware Arabic LLM
by: Alnumay, Yazeed, et al.
Published: (2025)
by: Alnumay, Yazeed, et al.
Published: (2025)
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
by: Elmadani, Khalid N., et al.
Published: (2025)
by: Elmadani, Khalid N., et al.
Published: (2025)
TSLFormer: A Lightweight Transformer Model for Turkish Sign Language Recognition Using Skeletal Landmarks
by: Ertürk, Kutay, et al.
Published: (2025)
by: Ertürk, Kutay, et al.
Published: (2025)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
by: Qian, Zhaozhi, et al.
Published: (2024)
by: Qian, Zhaozhi, et al.
Published: (2024)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
The Landscape of Arabic Large Language Models (ALLMs): A New Era for Arabic Language Technology
by: Al-Khalifa, Shahad, et al.
Published: (2025)
by: Al-Khalifa, Shahad, et al.
Published: (2025)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
by: Elesedy, Hayder, et al.
Published: (2024)
by: Elesedy, Hayder, et al.
Published: (2024)
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
by: Lekssays, Ahmed, et al.
Published: (2025)
by: Lekssays, Ahmed, et al.
Published: (2025)
UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages
by: Abdullahi, Tassallah, et al.
Published: (2026)
by: Abdullahi, Tassallah, et al.
Published: (2026)
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
by: Wróbel, Krzysztof, et al.
Published: (2026)
by: Wróbel, Krzysztof, et al.
Published: (2026)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
by: Habash, Nizar, et al.
Published: (2024)
by: Habash, Nizar, et al.
Published: (2024)
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
by: Koto, Fajri, et al.
Published: (2024)
by: Koto, Fajri, et al.
Published: (2024)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
by: Tasawong, Panuthep, et al.
Published: (2026)
by: Tasawong, Panuthep, et al.
Published: (2026)
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention
by: Jha, Prince, et al.
Published: (2024)
by: Jha, Prince, et al.
Published: (2024)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
by: Alhumud, Anas, et al.
Published: (2026)
by: Alhumud, Anas, et al.
Published: (2026)
Arabic Large Language Models for Medical Text Generation
by: Allam, Abdulrahman, et al.
Published: (2025)
by: Allam, Abdulrahman, et al.
Published: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
AlcLaM: Arabic Dialectal Language Model
by: Ahmed, Murtadha, et al.
Published: (2024)
by: Ahmed, Murtadha, et al.
Published: (2024)
Similar Items
-
PAM: Training Policy-Aligned Moderation Filters at Scale
by: Fatehkia, Masoomali, et al.
Published: (2025) -
Tool Calling for Arabic LLMs: Data Strategies and Instruction Tuning
by: Ersoy, Asim, et al.
Published: (2025) -
Do I Really Know? Learning Factual Self-Verification for Hallucination Reduction
by: Altinisik, Enes, et al.
Published: (2026) -
Explaining the role of Intrinsic Dimensionality in Adversarial Training
by: Altinisik, Enes, et al.
Published: (2024) -
Fanar 2.0: Arabic Generative AI Stack
by: FANAR TEAM, et al.
Published: (2026)