Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | AlQadi, Leen, Alzubaidi, Ahmed, Alyafeai, Mohammed, Alobeidli, Hamza, Alhammadi, Maitha, Alsuwaidi, Shaikha, Alkaabi, Omar, Boussaha, Basma El Amel, Hacid, Hakim |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
by: Alzubaidi, Ahmed, et al.
Published: (2025)
by: Alzubaidi, Ahmed, et al.
Published: (2025)
3LM: Bridging Arabic, STEM, and Code through Benchmarking
by: Boussaha, Basma El Amel, et al.
Published: (2025)
by: Boussaha, Basma El Amel, et al.
Published: (2025)
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024)
by: Malartic, Quentin, et al.
Published: (2024)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
by: Firdoussi, Aymane El, et al.
Published: (2024)
by: Firdoussi, Aymane El, et al.
Published: (2024)
New Variations and Structural Refinements of Discrete Weighted Jensen and Hermite–Hadamard Inequalities Using ( α , m)‐Convex Mappings
by: Shama Firdous, et al.
Published: (2026)
by: Shama Firdous, et al.
Published: (2026)
Adaptive CUSUM Control Chart With Variable Sample Size for Monitoring Under Measurement Error
by: Abdullah Ali H. Ahmadini, et al.
Published: (2026)
by: Abdullah Ali H. Ahmadini, et al.
Published: (2026)
Alignment with Preference Optimization Is All You Need for LLM Safety
by: Alami, Reda, et al.
Published: (2024)
by: Alami, Reda, et al.
Published: (2024)
Poem Meter Classification of Recited Arabic Poetry: Integrating High-Resource Systems for a Low-Resource Task
by: Al-Shaibani, Maged S., et al.
Published: (2025)
by: Al-Shaibani, Maged S., et al.
Published: (2025)
Re-thinking Human Activity Recognition with Hierarchy-aware Label Relationship Modeling
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
ArabicNumBench: Evaluating Arabic Number Reading in Large Language Models
by: Alhumud, Anas, et al.
Published: (2026)
by: Alhumud, Anas, et al.
Published: (2026)
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models
by: Yagoubi, Mouadh, et al.
Published: (2025)
by: Yagoubi, Mouadh, et al.
Published: (2025)
PORT: Preference Optimization on Reasoning Traces
by: Lahlou, Salem, et al.
Published: (2024)
by: Lahlou, Salem, et al.
Published: (2024)
The Increasing Importance of Breast Cancer in the United Arab Emirates
by: Ruqaiyyah Siddiqui, et al.
Published: (2025)
by: Ruqaiyyah Siddiqui, et al.
Published: (2025)
Data Quality in Edge Machine Learning: A State-of-the-Art Survey
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2024)
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2024)
Perceptions of Nursing Faculty on Utilizing AI Tools in Academic Writing and Publication Productivity: A Cross‐Sectional Study
by: Maitha Al Salti, et al.
Published: (2025)
by: Maitha Al Salti, et al.
Published: (2025)
Arab Americans in the United States
by: Al-Kuwari, Shaikha H.
Published: (2024)
by: Al-Kuwari, Shaikha H.
Published: (2024)
Cutting-Edge Innovations in Teaching, Leadership, Technology, and Assessment
by: Abdallah, Asma Khaleel, et al.
Published: (2024)
by: Abdallah, Asma Khaleel, et al.
Published: (2024)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
From Language to Action in Arabic: Reliable Structured Tool Calling via Data-Centric Fine-Tuning
by: Nacar, Omer, et al.
Published: (2026)
by: Nacar, Omer, et al.
Published: (2026)
Ionospheric Scintillation Forecasting Using Machine Learning
by: Halawa, Sultan, et al.
Published: (2024)
by: Halawa, Sultan, et al.
Published: (2024)
Evaluating the Effects of Arabic‐Medium Behavioral Skills Training on Parenting Program Facilitators in the United Arab Emirates
by: Amina Maliki, et al.
Published: (2025)
by: Amina Maliki, et al.
Published: (2025)
Chapter 7 Arabic mobile game localizations
by: Al‑Batineh, Mohammed
Published: (2025)
by: Al‑Batineh, Mohammed
Published: (2025)
MAGNETO: Edge AI for Human Activity Recognition -- Privacy and Personalization
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
ALRM: Agentic LLM for Robotic Manipulation
by: Santos, Vitor Gaboardi dos, et al.
Published: (2026)
by: Santos, Vitor Gaboardi dos, et al.
Published: (2026)
Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
by: Belgoumri, Mohammed Djameleddine, et al.
Published: (2025)
Mapping the Landscape of Generative Language Models in Dental Education: A Comparison Between ChatGPT and Google Bard
by: Shaikha Aldukhail
Published: (2024)
by: Shaikha Aldukhail
Published: (2024)
Development of a full-Scale approach to predict overlay reflective crack
by: Zhu, Zehui, et al.
Published: (2024)
by: Zhu, Zehui, et al.
Published: (2024)
SIFT-Aided Rectified 2D-DIC for Displacement and Strain Measurements in Asphalt Concrete Testing
by: Zhu, Zehui, et al.
Published: (2024)
by: Zhu, Zehui, et al.
Published: (2024)
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
by: Zuo, Jingwei, et al.
Published: (2025)
by: Zuo, Jingwei, et al.
Published: (2025)
Reliability of 1D radiative-convective photochemical-equilibrium retrievals on transit spectra of WASP-107b
by: Konings, Thomas, et al.
Published: (2025)
by: Konings, Thomas, et al.
Published: (2025)
MeXtract: Light-Weight Metadata Extraction from Scientific Papers
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
MOLE: Metadata Extraction and Validation in Scientific Papers Using LLMs
by: Alyafeai, Zaid, et al.
Published: (2025)
by: Alyafeai, Zaid, et al.
Published: (2025)
AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering
by: Alhuzali, Hassan, et al.
Published: (2025)
by: Alhuzali, Hassan, et al.
Published: (2025)
Falcon Mamba: The First Competitive Attention-free 7B Language Model
by: Zuo, Jingwei, et al.
Published: (2024)
by: Zuo, Jingwei, et al.
Published: (2024)
PSSF: Early osteoarthritis detection using physical synthetic knee X-ray scans and AI radiomics models
by: Alzubaidi, Abbas, et al.
Published: (2026)
by: Alzubaidi, Abbas, et al.
Published: (2026)
Stressors, Needs and Satisfaction of Families of Critically Ill Arabic Patients: A Systematic Review
by: Khaled Mohammed Al‐Sayaghi, et al.
Published: (2025)
by: Khaled Mohammed Al‐Sayaghi, et al.
Published: (2025)
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
by: Alyafeai, Zaid, et al.
Published: (2024)
by: Alyafeai, Zaid, et al.
Published: (2024)
Seismic Performance of Low‐Damage Prefabricated Segmental Post‐Tensioned Bridge Columns With Hybrid Connections
by: Tariq Al‐Qadi, et al.
Published: (2026)
by: Tariq Al‐Qadi, et al.
Published: (2026)
Comparative Efficacy of Photodynamic Therapy Versus Cryotherapy for Actinic Keratosis: A Systematic Review and Meta‐Analysis of Randomized Controlled Trials
by: Saleh Aldraibi, et al.
Published: (2026)
by: Saleh Aldraibi, et al.
Published: (2026)
Similar Items
-
Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps
by: Alzubaidi, Ahmed, et al.
Published: (2025) -
3LM: Bridging Arabic, STEM, and Code through Benchmarking
by: Boussaha, Basma El Amel, et al.
Published: (2025) -
Falcon2-11B Technical Report
by: Malartic, Quentin, et al.
Published: (2024) -
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
by: Firdoussi, Aymane El, et al.
Published: (2024) -
New Variations and Structural Refinements of Discrete Weighted Jensen and Hermite–Hadamard Inequalities Using ( α , m)‐Convex Mappings
by: Shama Firdous, et al.
Published: (2026)