NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
Fuente:
arXiv
Saved in:
| Main Authors: | Mekki, Abdellah El, Atou, Houdaifa, Nacar, Omer, Shehata, Shady, Abdul-Mageed, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs
by: Mekki, Abdellah El, et al.
Published: (2024)
by: Mekki, Abdellah El, et al.
Published: (2024)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs
by: Saeed, Muhammed, et al.
Published: (2024)
by: Saeed, Muhammed, et al.
Published: (2024)
Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
by: Saeed, Muhammed, et al.
Published: (2025)
by: Saeed, Muhammed, et al.
Published: (2025)
UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat
by: Nacar, Omer
Published: (2025)
by: Nacar, Omer
Published: (2025)
Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs
by: Mekki, Abdellah El, et al.
Published: (2026)
by: Mekki, Abdellah El, et al.
Published: (2026)
GemMaroc: Unlocking Darija Proficiency in LLMs with Minimal Data
by: Skiredj, Abderrahman, et al.
Published: (2025)
by: Skiredj, Abderrahman, et al.
Published: (2025)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Beyond Content: How Grammatical Gender Shapes Visual Representation in Text-to-Image Models
by: Saeed, Muhammed, et al.
Published: (2025)
by: Saeed, Muhammed, et al.
Published: (2025)
Zero-Shot Context-Aware ASR for Diverse Arabic Varieties
by: Talafha, Bashar, et al.
Published: (2025)
by: Talafha, Bashar, et al.
Published: (2025)
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
by: Magdy, Samar M., et al.
Published: (2025)
by: Magdy, Samar M., et al.
Published: (2025)
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
by: Kwon, Sang Yun, et al.
Published: (2026)
by: Kwon, Sang Yun, et al.
Published: (2026)
Enhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning
by: Nacar, Omer, et al.
Published: (2024)
by: Nacar, Omer, et al.
Published: (2024)
Dallah: A Dialect-Aware Multimodal Large Language Model for Arabic
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
by: Doan, Khai Duy, et al.
Published: (2024)
by: Doan, Khai Duy, et al.
Published: (2024)
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
by: Ibrahim, Adham, et al.
Published: (2024)
by: Ibrahim, Adham, et al.
Published: (2024)
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
by: Wu, Minghao, et al.
Published: (2023)
by: Wu, Minghao, et al.
Published: (2023)
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability
by: Chen, Wei-Rui, et al.
Published: (2023)
by: Chen, Wei-Rui, et al.
Published: (2023)
Detecting Propaganda Techniques in Code-Switched Social Media Text
by: Salman, Muhammad Umar, et al.
Published: (2023)
by: Salman, Muhammad Umar, et al.
Published: (2023)
Arabic Automatic Story Generation with Large Language Models
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
by: El-Shangiti, Ahmed Oumar, et al.
Published: (2024)
ASR Under Noise: Exploring Robustness for Sundanese and Javanese
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
by: Pranida, Salsabila Zahirah, et al.
Published: (2025)
To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
On Barriers to Archival Audio Processing
by: Sullivan, Peter, et al.
Published: (2025)
by: Sullivan, Peter, et al.
Published: (2025)
Toucan: Many-to-Many Translation for 150 African Language Pairs
by: Elmadany, AbdelRahim, et al.
Published: (2024)
by: Elmadany, AbdelRahim, et al.
Published: (2024)
Cheetah: Natural Language Generation for 517 African Languages
by: Adebara, Ife, et al.
Published: (2024)
by: Adebara, Ife, et al.
Published: (2024)
NurseLLM: The First Specialized Language Model for Nursing
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2025)
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Distilling Text Style Transfer With Self-Explanation From LLMs
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
ArFake: A Multi-Dialect Benchmark and Baselines for Arabic Spoof-Speech Detection
by: Maged, Mohamed, et al.
Published: (2025)
by: Maged, Mohamed, et al.
Published: (2025)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Interplay of Machine Translation, Diacritics, and Diacritization
by: Chen, Wei-Rui, et al.
Published: (2024)
by: Chen, Wei-Rui, et al.
Published: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
by: Waheed, Abdul, et al.
Published: (2024)
by: Waheed, Abdul, et al.
Published: (2024)
ArabEmoNet: A Lightweight Hybrid 2D CNN-BiLSTM Model with Attention for Robust Arabic Speech Emotion Recognition
by: Abouzeid, Ali, et al.
Published: (2025)
by: Abouzeid, Ali, et al.
Published: (2025)
LLM Performance Predictors are good initializers for Architecture Search
by: Jawahar, Ganesh, et al.
Published: (2023)
by: Jawahar, Ganesh, et al.
Published: (2023)
Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks
by: Alwajih, Fakhraddin, et al.
Published: (2024)
by: Alwajih, Fakhraddin, et al.
Published: (2024)
Similar Items
-
Effective Self-Mining of In-Context Examples for Unsupervised Machine Translation with LLMs
by: Mekki, Abdellah El, et al.
Published: (2024) -
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025) -
Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks
by: Bhatia, Gagan, et al.
Published: (2024) -
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
by: Alwajih, Fakhraddin, et al.
Published: (2025) -
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)