AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Hasanaath, Ahmed, Alansari, Aisha, Ashraf, Ahmed, Salmane, Chafik, Luqman, Hamzah, Ezzini, Saad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija
by: Chafik, Salmane, et al.
Published: (2025)
by: Chafik, Salmane, et al.
Published: (2025)
Large Language Models Hallucination: A Comprehensive Survey
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)
by: Alansari, Aisha, et al.
Published: (2026)
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
by: Alharbi, Maram, et al.
Published: (2025)
by: Alharbi, Maram, et al.
Published: (2025)
AraFinNLP 2024: The First Arabic Financial NLP Shared Task
by: Malaysha, Sanad, et al.
Published: (2024)
by: Malaysha, Sanad, et al.
Published: (2024)
USTM: Unified Spatial and Temporal Modeling for Continuous Sign Language Recognition
by: Hasanaath, Ahmed Abul, et al.
Published: (2025)
by: Hasanaath, Ahmed Abul, et al.
Published: (2025)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
LeGo-Code: Can Modular Curriculum Learning Advance Complex Code Generation? Insights from Text-to-SQL
by: Chafik, Salmane, et al.
Published: (2026)
by: Chafik, Salmane, et al.
Published: (2026)
FSBI: Deepfakes Detection with Frequency Enhanced Self-Blended Images
by: Hasanaath, Ahmed Abul, et al.
Published: (2024)
by: Hasanaath, Ahmed Abul, et al.
Published: (2024)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
by: Alshaikh, Rana, et al.
Published: (2025)
by: Alshaikh, Rana, et al.
Published: (2025)
M-DAIGT: A Shared Task on Multi-Domain Detection of AI-Generated Text
by: Lamsiyah, Salima, et al.
Published: (2025)
by: Lamsiyah, Salima, et al.
Published: (2025)
AraSpider: Democratizing Arabic-to-SQL
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
DarijaBanking: A New Resource for Overcoming Language Barriers in Banking Intent Detection for Moroccan Arabic Speakers
by: Skiredj, Abderrahman, et al.
Published: (2024)
by: Skiredj, Abderrahman, et al.
Published: (2024)
A Comparative Study of Continuous Sign Language Recognition Techniques
by: Alyami, Sarah, et al.
Published: (2024)
by: Alyami, Sarah, et al.
Published: (2024)
AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
by: El-Haj, Mo, et al.
Published: (2025)
by: El-Haj, Mo, et al.
Published: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
!MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering through Prompt Engineering and Ensemble Learning
by: Tarek, Mohamed, et al.
Published: (2025)
by: Tarek, Mohamed, et al.
Published: (2025)
AraS2P: Arabic Speech-to-Phonemes System
by: Matar, Bassam, et al.
Published: (2025)
by: Matar, Bassam, et al.
Published: (2025)
Ara-Best-RQ: Multi Dialectal Arabic SSL
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
AraSpot: Arabic Spoken Command Spotting
by: Salhab, Mahmoud, et al.
Published: (2023)
by: Salhab, Mahmoud, et al.
Published: (2023)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
by: Alwajih, Fakhraddin, et al.
Published: (2025)
by: Alwajih, Fakhraddin, et al.
Published: (2025)
Creating Arabic LLM Prompts at Scale
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
by: El-Sheikh, Abdelrahman, et al.
Published: (2024)
QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning
by: AL-Smadi, Mohammad
Published: (2026)
by: AL-Smadi, Mohammad
Published: (2026)
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
by: Abdelnabi, Sahar, et al.
Published: (2025)
by: Abdelnabi, Sahar, et al.
Published: (2025)
Arabic Dataset for LLM Safeguard Evaluation
by: Ashraf, Yasser, et al.
Published: (2024)
by: Ashraf, Yasser, et al.
Published: (2024)
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
by: Salhab, Mahmoud, et al.
Published: (2024)
by: Salhab, Mahmoud, et al.
Published: (2024)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering
by: Alhuzali, Hassan, et al.
Published: (2025)
by: Alhuzali, Hassan, et al.
Published: (2025)
MedAraBench: Large-Scale Arabic Medical Question Answering Dataset and Benchmark
by: Abu-Daoud, Mouath, et al.
Published: (2026)
by: Abu-Daoud, Mouath, et al.
Published: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
by: Staniek, Michael, et al.
Published: (2025)
by: Staniek, Michael, et al.
Published: (2025)
Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning
by: Chan, Jason, et al.
Published: (2025)
by: Chan, Jason, et al.
Published: (2025)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
by: Qarah, Faisal
Published: (2024)
by: Qarah, Faisal
Published: (2024)
AraModernBERT: Transtokenized Initialization and Long-Context Encoder Modeling for Arabic
by: Elshehy, Omar, et al.
Published: (2026)
by: Elshehy, Omar, et al.
Published: (2026)
dzFinNlp at AraFinNLP: Improving Intent Detection in Financial Conversational Agents
by: Lichouri, Mohamed, et al.
Published: (2024)
by: Lichouri, Mohamed, et al.
Published: (2024)
Revisiting Common Assumptions about Arabic Dialects in NLP
by: Keleg, Amr, et al.
Published: (2025)
by: Keleg, Amr, et al.
Published: (2025)
AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse
by: Sharqawi, Esra'a, et al.
Published: (2026)
by: Sharqawi, Esra'a, et al.
Published: (2026)
Similar Items
-
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
by: Alansari, Aisha, et al.
Published: (2025) -
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
by: Alansari, Aisha, et al.
Published: (2025) -
Dialect2SQL: A Novel Text-to-SQL Dataset for Arabic Dialects with a Focus on Moroccan Darija
by: Chafik, Salmane, et al.
Published: (2025) -
Large Language Models Hallucination: A Comprehensive Survey
by: Alansari, Aisha, et al.
Published: (2025) -
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)