AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Alansari, Aisha, Luqman, Hamzah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)
by: Alansari, Aisha, et al.
Published: (2026)
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025)
by: Hasanaath, Ahmed, et al.
Published: (2025)
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
Large Language Models Hallucination: A Comprehensive Survey
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)
by: Alghamdi, Emad A., et al.
Published: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
by: Zhao, Yaxin, et al.
Published: (2025)
by: Zhao, Yaxin, et al.
Published: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
A Comparative Study of Continuous Sign Language Recognition Techniques
by: Alyami, Sarah, et al.
Published: (2024)
by: Alyami, Sarah, et al.
Published: (2024)
AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
by: El-Haj, Mo, et al.
Published: (2025)
by: El-Haj, Mo, et al.
Published: (2025)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
by: Niazi, Ruhallah, et al.
Published: (2026)
by: Niazi, Ruhallah, et al.
Published: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
AraTable: Benchmarking LLMs' Reasoning and Understanding of Arabic Tabular Data
by: Alshaikh, Rana, et al.
Published: (2025)
by: Alshaikh, Rana, et al.
Published: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
by: Abdallah, Mohamed A., et al.
Published: (2025)
by: Abdallah, Mohamed A., et al.
Published: (2025)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
Ara-HOPE: Human-Centric Post-Editing Evaluation for Dialectal Arabic to Modern Standard Arabic Translation
by: Alabdullah, Abdullah, et al.
Published: (2025)
by: Alabdullah, Abdullah, et al.
Published: (2025)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
by: Luo, Wen, et al.
Published: (2024)
by: Luo, Wen, et al.
Published: (2024)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025)
by: Dasgupta, Sharanya, et al.
Published: (2025)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
by: Mustapha, Ahmad, et al.
Published: (2024)
by: Mustapha, Ahmad, et al.
Published: (2024)
HalluZig: Hallucination Detection using Zigzag Persistence
by: Samaga, Shreyas N., et al.
Published: (2026)
by: Samaga, Shreyas N., et al.
Published: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
by: Li, Tianyi, et al.
Published: (2024)
by: Li, Tianyi, et al.
Published: (2024)
AraS2P: Arabic Speech-to-Phonemes System
by: Matar, Bassam, et al.
Published: (2025)
by: Matar, Bassam, et al.
Published: (2025)
Ara-Best-RQ: Multi Dialectal Arabic SSL
by: Elleuch, Haroun, et al.
Published: (2026)
by: Elleuch, Haroun, et al.
Published: (2026)
AraSpot: Arabic Spoken Command Spotting
by: Salhab, Mahmoud, et al.
Published: (2023)
by: Salhab, Mahmoud, et al.
Published: (2023)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
by: Urlana, Ashok, et al.
Published: (2025)
by: Urlana, Ashok, et al.
Published: (2025)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
by: Yeh, Min-Hsuan, et al.
Published: (2025)
by: Yeh, Min-Hsuan, et al.
Published: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
BUSTED at AraGenEval Shared Task: A Comparative Study of Transformer-Based Models for Arabic AI-Generated Text Detection
by: Zain, Ali, et al.
Published: (2025)
by: Zain, Ali, et al.
Published: (2025)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
AraSpider: Democratizing Arabic-to-SQL
by: Heakl, Ahmed, et al.
Published: (2024)
by: Heakl, Ahmed, et al.
Published: (2024)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
by: Salhab, Mahmoud, et al.
Published: (2024)
by: Salhab, Mahmoud, et al.
Published: (2024)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
by: Anaokar, Spandan, et al.
Published: (2025)
by: Anaokar, Spandan, et al.
Published: (2025)
AraFinNLP 2024: The First Arabic Financial NLP Shared Task
by: Malaysha, Sanad, et al.
Published: (2024)
by: Malaysha, Sanad, et al.
Published: (2024)
A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment
by: Elmadani, Khalid N., et al.
Published: (2025)
by: Elmadani, Khalid N., et al.
Published: (2025)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
by: Habash, Nizar, et al.
Published: (2024)
by: Habash, Nizar, et al.
Published: (2024)
Similar Items
-
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026) -
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP
by: Hasanaath, Ahmed, et al.
Published: (2025) -
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
by: Alansari, Aisha, et al.
Published: (2025) -
Large Language Models Hallucination: A Comprehensive Survey
by: Alansari, Aisha, et al.
Published: (2025) -
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic
by: Alghamdi, Emad A., et al.
Published: (2024)