AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zbeeb, Mohammad, Hammoud, Hasan Abed Al Kader, Mukalled, Sina, Rizk, Nadine, Karnib, Fatima, Lakkis, Issam, Mohanna, Ammar, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
TAPS: Task Aware Proposal Distributions for Speculative Sampling
by: Zbib, Mohamad, et al.
Published: (2026)
by: Zbib, Mohamad, et al.
Published: (2026)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation
by: Slim, Ali, et al.
Published: (2026)
by: Slim, Ali, et al.
Published: (2026)
On the Importance of Pretraining Data Alignment for Atomic Property Prediction
by: Ghunaim, Yasir, et al.
Published: (2025)
by: Ghunaim, Yasir, et al.
Published: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
An Embarrassingly Simple Defense Against LLM Abliteration Attacks
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
by: Shairah, Harethah Abu, et al.
Published: (2025)
by: Shairah, Harethah Abu, et al.
Published: (2025)
Optimizing Deep Neural Networks using Safety-Guided Self Compression
by: Zbeeb, Mohammad, et al.
Published: (2025)
by: Zbeeb, Mohammad, et al.
Published: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Forget Less, Retain More: A Lightweight Regularizer for Rehearsal-Based Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
by: Alssum, Lama, et al.
Published: (2025)
by: Alssum, Lama, et al.
Published: (2025)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
by: Prabhu, Ameya, et al.
Published: (2023)
by: Prabhu, Ameya, et al.
Published: (2023)
Reinforcement Learning, Optimal Control, and Bayesian Filtering in Data Assimilation
by: Hammoud, Abed
Published: (2026)
by: Hammoud, Abed
Published: (2026)
Sectoral‐Driven Framework for Dynamic Water Stress Assessment and Management
by: Ali Karnib
Published: (2025)
by: Ali Karnib
Published: (2025)
DamascusTeam at NLP4IF2021: Fighting the Arabic COVID-19 Infodemic on Twitter Using AraBERT
by: Ahmad Hussein, et al.
Published: (2021)
by: Ahmad Hussein, et al.
Published: (2021)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
AraSpider: Democratizing Arabic-to-SQL
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
AraSpot: Arabic Spoken Command Spotting
by: Salhab, Mahmoud, et al.
Published: (2023)
by: Salhab, Mahmoud, et al.
Published: (2023)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
Randomized Asymmetric Chain of LoRA: The First Meaningful Theoretical Framework for Low-Rank Adaptation
by: Malinovsky, Grigory, et al.
Published: (2024)
by: Malinovsky, Grigory, et al.
Published: (2024)
AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse
by: Sharqawi, Esra'a, et al.
Published: (2026)
by: Sharqawi, Esra'a, et al.
Published: (2026)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Ara-Best-RQ: Multi Dialectal Arabic SSL
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
AraS2P: Arabic Speech-to-Phonemes System
by: Matar, Bassam, et al.
Published: (2025)
by: Matar, Bassam, et al.
Published: (2025)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
by: Yang, Yibo, et al.
Published: (2024)
by: Yang, Yibo, et al.
Published: (2024)
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
by: Mousi, Basel, et al.
Published: (2024)
by: Mousi, Basel, et al.
Published: (2024)
LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
by: Lian, Da-Chen, et al.
Published: (2025)
by: Lian, Da-Chen, et al.
Published: (2025)
EmoAra: Emotion-Preserving English Speech Transcription and Cross-Lingual Translation with Arabic Text-to-Speech
by: Hassan, Besher, et al.
Published: (2026)
by: Hassan, Besher, et al.
Published: (2026)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
Chained Prompting for Better Systematic Review Search Strategies
by: Nasser, Fatima, et al.
Published: (2025)
by: Nasser, Fatima, et al.
Published: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
by: Alshaikh, Rana, et al.
Published: (2025)
by: Alshaikh, Rana, et al.
Published: (2025)
AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
by: El-Haj, Mo, et al.
Published: (2025)
by: El-Haj, Mo, et al.
Published: (2025)
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
by: Salhab, Mahmoud, et al.
Published: (2024)
by: Salhab, Mahmoud, et al.
Published: (2024)
Similar Items
-
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025) -
Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
by: Zbeeb, Mohammad, et al.
Published: (2025) -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
by: Zbib, Mohamad, et al.
Published: (2026) -
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025) -
QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation
by: Slim, Ali, et al.
Published: (2026)