PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
Fuente:
arXiv
Saved in:
| Main Authors: | Hussain, Khizar, Kantarcioglu, Murat |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
by: Hussain, Khizar, et al.
Published: (2026)
by: Hussain, Khizar, et al.
Published: (2026)
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024)
by: Luo, Junliang, et al.
Published: (2024)
Budget-Aware Routing for Long Clinical Text
by: Qureshi, Khizar, et al.
Published: (2026)
by: Qureshi, Khizar, et al.
Published: (2026)
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
by: Basit, Abdul, et al.
Published: (2024)
by: Basit, Abdul, et al.
Published: (2024)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
by: Atasoy, I. F., et al.
Published: (2026)
by: Atasoy, I. F., et al.
Published: (2026)
Hallucination Detection with Small Language Models
by: Cheung, Ming
Published: (2025)
by: Cheung, Ming
Published: (2025)
Hallucination Detection with the Internal Layers of LLMs
by: Preiß, Martin
Published: (2025)
by: Preiß, Martin
Published: (2025)
Towards Long Context Hallucination Detection
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
How to Detect and Defeat Molecular Mirage: A Metric-Driven Benchmark for Hallucination in LLM-based Molecular Comprehension
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
LettuceDetect: A Hallucination Detection Framework for RAG Applications
by: Kovács, Ádám, et al.
Published: (2025)
by: Kovács, Ádám, et al.
Published: (2025)
Sanity Checks for Long-Form Hallucination Detection
by: Zollicoffer, Geigh, et al.
Published: (2026)
by: Zollicoffer, Geigh, et al.
Published: (2026)
Enhancing Hallucination Detection via Future Context
by: Lee, Joosung, et al.
Published: (2025)
by: Lee, Joosung, et al.
Published: (2025)
Comparing Hallucination Detection Metrics for Multilingual Generation
by: Kang, Haoqiang, et al.
Published: (2024)
by: Kang, Haoqiang, et al.
Published: (2024)
On Early Detection of Hallucinations in Factual Question Answering
by: Snyder, Ben, et al.
Published: (2023)
by: Snyder, Ben, et al.
Published: (2023)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
by: Srey, Ponhvoan, et al.
Published: (2025)
by: Srey, Ponhvoan, et al.
Published: (2025)
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation
by: Sun, Siqi, et al.
Published: (2026)
by: Sun, Siqi, et al.
Published: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
by: Liu, Litian, et al.
Published: (2026)
by: Liu, Litian, et al.
Published: (2026)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
by: Ghanem, Hussam, et al.
Published: (2025)
by: Ghanem, Hussam, et al.
Published: (2025)
Hallucination Detection-Guided Preference Optimization for Clinical Summarization
by: Seethakantha, Shamanth Kuthpadi, et al.
Published: (2026)
by: Seethakantha, Shamanth Kuthpadi, et al.
Published: (2026)
Enhancing Uncertainty Modeling with Semantic Graph for Hallucination Detection
by: Chen, Kedi, et al.
Published: (2025)
by: Chen, Kedi, et al.
Published: (2025)
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
VeriTrail: Closed-Domain Hallucination Detection with Traceability
by: Metropolitansky, Dasha, et al.
Published: (2025)
by: Metropolitansky, Dasha, et al.
Published: (2025)
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
by: Bazarova, Alexandra, et al.
Published: (2025)
by: Bazarova, Alexandra, et al.
Published: (2025)
HARP: Hallucination Detection via Reasoning Subspace Projection
by: Hu, Junjie, et al.
Published: (2025)
by: Hu, Junjie, et al.
Published: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Turk-LettuceDetect: A Hallucination Detection Models for Turkish RAG Applications
by: Taş, Selva, et al.
Published: (2025)
by: Taş, Selva, et al.
Published: (2025)
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
by: Gautam, Arpit Singh, et al.
Published: (2026)
by: Gautam, Arpit Singh, et al.
Published: (2026)
The First Token Knows: Single-Decode Confidence for Hallucination Detection
by: Gabriel, Mina
Published: (2026)
by: Gabriel, Mina
Published: (2026)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
by: Gogoulou, Evangelia, et al.
Published: (2025)
by: Gogoulou, Evangelia, et al.
Published: (2025)
Ask a Local: Detecting Hallucinations With Specialized Model Divergence
by: Creo, Aldan, et al.
Published: (2025)
by: Creo, Aldan, et al.
Published: (2025)
Trustworthy AI for Medicine: Continuous Hallucination Detection and Elimination with CHECK
by: Garcia-Fernandez, Carlos, et al.
Published: (2025)
by: Garcia-Fernandez, Carlos, et al.
Published: (2025)
Similar Items
-
Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses
by: Hussain, Khizar, et al.
Published: (2026) -
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024) -
Budget-Aware Routing for Long Clinical Text
by: Qureshi, Khizar, et al.
Published: (2026) -
MedAide: Leveraging Large Language Models for On-Premise Medical Assistance on Edge Devices
by: Basit, Abdul, et al.
Published: (2024) -
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)