HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Emmy, Gangal, Varun, Yu, Michael, Tao, Zhuofu, Singh, Karan, Kumar, Sachin, Feng, Steven Y. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Unified Definition of Hallucination: It's The World Model, Stupid!
di: Liu, Emmy, et al.
Pubblicazione: (2025)
di: Liu, Emmy, et al.
Pubblicazione: (2025)
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
di: Singh, Karan, et al.
Pubblicazione: (2026)
di: Singh, Karan, et al.
Pubblicazione: (2026)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
di: Emery, Deanna, et al.
Pubblicazione: (2025)
di: Emery, Deanna, et al.
Pubblicazione: (2025)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
di: Urlana, Ashok, et al.
Pubblicazione: (2025)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
di: Alansari, Aisha, et al.
Pubblicazione: (2026)
di: Alansari, Aisha, et al.
Pubblicazione: (2026)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
di: Liu, Xuannan, et al.
Pubblicazione: (2026)
di: Liu, Xuannan, et al.
Pubblicazione: (2026)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
di: Yeh, Min-Hsuan, et al.
Pubblicazione: (2025)
di: Yeh, Min-Hsuan, et al.
Pubblicazione: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
di: Hosseini, Mohammad, et al.
Pubblicazione: (2025)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
di: Fan, Dongyang, et al.
Pubblicazione: (2026)
Enhancing Hallucination Detection through Perturbation-Based Synthetic Data Generation in System Responses
di: Zhang, Dongxu, et al.
Pubblicazione: (2024)
di: Zhang, Dongxu, et al.
Pubblicazione: (2024)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
di: Anaokar, Spandan, et al.
Pubblicazione: (2025)
di: Anaokar, Spandan, et al.
Pubblicazione: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
di: Zhao, Yaxin, et al.
Pubblicazione: (2025)
di: Zhao, Yaxin, et al.
Pubblicazione: (2025)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
di: Kumar, Mahesh, et al.
Pubblicazione: (2026)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
di: Hosseini, Faezeh, et al.
Pubblicazione: (2026)
di: Hosseini, Faezeh, et al.
Pubblicazione: (2026)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
di: Luo, Wen, et al.
Pubblicazione: (2024)
di: Luo, Wen, et al.
Pubblicazione: (2024)
HalluZig: Hallucination Detection using Zigzag Persistence
di: Samaga, Shreyas N., et al.
Pubblicazione: (2026)
di: Samaga, Shreyas N., et al.
Pubblicazione: (2026)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
di: Chen, Boshui, et al.
Pubblicazione: (2026)
di: Chen, Boshui, et al.
Pubblicazione: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
di: Li, Tianyi, et al.
Pubblicazione: (2024)
di: Li, Tianyi, et al.
Pubblicazione: (2024)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
di: Cherif, Ahmed
Pubblicazione: (2026)
di: Cherif, Ahmed
Pubblicazione: (2026)
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
di: Bergeron, Loris, et al.
Pubblicazione: (2025)
di: Bergeron, Loris, et al.
Pubblicazione: (2025)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
di: Vashistha, Sachin, et al.
Pubblicazione: (2025)
di: Vashistha, Sachin, et al.
Pubblicazione: (2025)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
di: Dasgupta, Sharanya, et al.
Pubblicazione: (2025)
di: Dasgupta, Sharanya, et al.
Pubblicazione: (2025)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
di: Alansari, Aisha, et al.
Pubblicazione: (2025)
di: Alansari, Aisha, et al.
Pubblicazione: (2025)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
di: Noël, Valentin, et al.
Pubblicazione: (2025)
di: Noël, Valentin, et al.
Pubblicazione: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
Sparse Rewards Can Self-Train Dialogue Agents
di: Lattimer, Barrett Martin, et al.
Pubblicazione: (2024)
di: Lattimer, Barrett Martin, et al.
Pubblicazione: (2024)
HALT: Hallucination Assessment via Log-probs as Time series
di: Shapiro, Ahmad, et al.
Pubblicazione: (2026)
di: Shapiro, Ahmad, et al.
Pubblicazione: (2026)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
di: Ridder, Fabian, et al.
Pubblicazione: (2024)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
di: Nath, Sujoy, et al.
Pubblicazione: (2025)
di: Nath, Sujoy, et al.
Pubblicazione: (2025)
RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
di: Hu, Xiangkun, et al.
Pubblicazione: (2024)
di: Hu, Xiangkun, et al.
Pubblicazione: (2024)
Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification
di: Sahu, Pritish, et al.
Pubblicazione: (2024)
di: Sahu, Pritish, et al.
Pubblicazione: (2024)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
di: Mao, Nathan, et al.
Pubblicazione: (2026)
di: Mao, Nathan, et al.
Pubblicazione: (2026)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
di: Meng, Jinxiang, et al.
Pubblicazione: (2026)
di: Meng, Jinxiang, et al.
Pubblicazione: (2026)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
di: Abdallah, Mohamed A., et al.
Pubblicazione: (2025)
di: Abdallah, Mohamed A., et al.
Pubblicazione: (2025)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
di: Sakai, Yusuke, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Unified Definition of Hallucination: It's The World Model, Stupid!
di: Liu, Emmy, et al.
Pubblicazione: (2025) -
To Memorize or to Retrieve: Scaling Laws for RAG-Considerate Pretraining
di: Singh, Karan, et al.
Pubblicazione: (2026) -
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
di: Emery, Deanna, et al.
Pubblicazione: (2025) -
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
di: Urlana, Ashok, et al.
Pubblicazione: (2025) -
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)