HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Wen, Shen, Tianshu, Li, Wei, Peng, Guangyue, Xuan, Richeng, Wang, Houfeng, Yang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
by: Yeh, Min-Hsuan, et al.
Published: (2025)
by: Yeh, Min-Hsuan, et al.
Published: (2025)
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
by: Luo, Wen, et al.
Published: (2025)
by: Luo, Wen, et al.
Published: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)
by: Alansari, Aisha, et al.
Published: (2026)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
ComperDial: Commonsense Persona-grounded Dialogue Dataset and Benchmark
by: Wakaki, Hiromi, et al.
Published: (2024)
by: Wakaki, Hiromi, et al.
Published: (2024)
Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
by: Anaokar, Spandan, et al.
Published: (2025)
by: Anaokar, Spandan, et al.
Published: (2025)
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models
by: Gao, Jian, et al.
Published: (2025)
by: Gao, Jian, et al.
Published: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
by: Hosseini, Faezeh, et al.
Published: (2026)
by: Hosseini, Faezeh, et al.
Published: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
by: Li, Tianyi, et al.
Published: (2024)
by: Li, Tianyi, et al.
Published: (2024)
CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
by: Shayanfar, Radin, et al.
Published: (2025)
by: Shayanfar, Radin, et al.
Published: (2025)
HalluZig: Hallucination Detection using Zigzag Persistence
by: Samaga, Shreyas N., et al.
Published: (2026)
by: Samaga, Shreyas N., et al.
Published: (2026)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026)
by: Liu, Emmy, et al.
Published: (2026)
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
by: Peng, Guangyue, et al.
Published: (2026)
by: Peng, Guangyue, et al.
Published: (2026)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
by: Urlana, Ashok, et al.
Published: (2025)
by: Urlana, Ashok, et al.
Published: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
by: Zhao, Yaxin, et al.
Published: (2025)
by: Zhao, Yaxin, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
by: Wei, Shouang, et al.
Published: (2025)
by: Wei, Shouang, et al.
Published: (2025)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
by: Chen, Kedi, et al.
Published: (2024)
by: Chen, Kedi, et al.
Published: (2024)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
by: Chen, Boshui, et al.
Published: (2026)
by: Chen, Boshui, et al.
Published: (2026)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
Detection-Correction Structure via General Language Model for Grammatical Error Correction
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
by: Kumar, Mahesh, et al.
Published: (2026)
by: Kumar, Mahesh, et al.
Published: (2026)
DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
by: Kim, Jiho, et al.
Published: (2024)
by: Kim, Jiho, et al.
Published: (2024)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
by: Chen, Xiangyan, et al.
Published: (2025)
by: Chen, Xiangyan, et al.
Published: (2025)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation
by: Xu, Ziyao, et al.
Published: (2024)
by: Xu, Ziyao, et al.
Published: (2024)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025)
by: Dasgupta, Sharanya, et al.
Published: (2025)
Similar Items
-
Explanation based In-Context Demonstrations Retrieval for Multilingual Grammatical Error Correction
by: Li, Wei, et al.
Published: (2025) -
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
by: Li, Wei, et al.
Published: (2025) -
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
by: Yeh, Min-Hsuan, et al.
Published: (2025) -
Odysseus Navigates the Sirens' Song: Dynamic Focus Decoding for Factual and Diverse Open-Ended Text Generation
by: Luo, Wen, et al.
Published: (2025) -
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)