On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Jing, Xiaonan, Billa, Srinivas, Godbout, Danny |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TravelBench : Exploring LLM Performance in Low-Resource Domains
by: Billa, Srinivas, et al.
Published: (2025)
by: Billa, Srinivas, et al.
Published: (2025)
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
The Geometric Anatomy of Capability Acquisition in Transformers
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
Quantifying Hallucinations in Language Language Models on Medical Textbooks
by: Colelough, Brandon C., et al.
Published: (2026)
by: Colelough, Brandon C., et al.
Published: (2026)
Supervisory Prompt Training
by: Billa, Jean Ghislain, et al.
Published: (2024)
by: Billa, Jean Ghislain, et al.
Published: (2024)
STORYSUMM: Evaluating Faithfulness in Story Summarization
by: Subbiah, Melanie, et al.
Published: (2024)
by: Subbiah, Melanie, et al.
Published: (2024)
Credal Transformer: A Principled Approach for Quantifying and Mitigating Hallucinations in Large Language Models
by: Ji, Shihao, et al.
Published: (2025)
by: Ji, Shihao, et al.
Published: (2025)
The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models
by: Müller, Heimo, et al.
Published: (2026)
by: Müller, Heimo, et al.
Published: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
by: Fayyaz, Mohsen, et al.
Published: (2024)
by: Fayyaz, Mohsen, et al.
Published: (2024)
One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
by: Wang, Pengbo, et al.
Published: (2025)
by: Wang, Pengbo, et al.
Published: (2025)
From RAG to Agentic RAG for Faithful Islamic Question Answering
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
by: Adlakha, Vaibhav, et al.
Published: (2023)
by: Adlakha, Vaibhav, et al.
Published: (2023)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
by: Rivera, Mauricio, et al.
Published: (2024)
by: Rivera, Mauricio, et al.
Published: (2024)
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
by: Wang, Haofeng, et al.
Published: (2025)
by: Wang, Haofeng, et al.
Published: (2025)
Lynx: An Open Source Hallucination Evaluation Model
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
by: Ravi, Selvan Sunitha, et al.
Published: (2024)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
by: Qi, Siya, et al.
Published: (2024)
by: Qi, Siya, et al.
Published: (2024)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
by: Ju, Li, et al.
Published: (2026)
by: Ju, Li, et al.
Published: (2026)
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
by: Hashem, Tahsina, et al.
Published: (2023)
by: Hashem, Tahsina, et al.
Published: (2023)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
by: Halperin, Igor
Published: (2025)
by: Halperin, Igor
Published: (2025)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
by: Yang, Nakyeong, et al.
Published: (2025)
by: Yang, Nakyeong, et al.
Published: (2025)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
by: Liu, Litian, et al.
Published: (2026)
by: Liu, Litian, et al.
Published: (2026)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
by: Yang, Zachary, et al.
Published: (2023)
by: Yang, Zachary, et al.
Published: (2023)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
by: Siegel, Noah Y., et al.
Published: (2024)
by: Siegel, Noah Y., et al.
Published: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
by: Alon, Bar, et al.
Published: (2026)
by: Alon, Bar, et al.
Published: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024)
by: Luo, Junliang, et al.
Published: (2024)
Retracing the Past: LLMs Emit Training Data When They Get Lost
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
by: Hao, Yijie, et al.
Published: (2025)
by: Hao, Yijie, et al.
Published: (2025)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
by: Doshi, Savan
Published: (2026)
by: Doshi, Savan
Published: (2026)
C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
by: Gu, Yuzhe, et al.
Published: (2024)
by: Gu, Yuzhe, et al.
Published: (2024)
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
by: Zhou, Yiqing, et al.
Published: (2025)
by: Zhou, Yiqing, et al.
Published: (2025)
Removal of Hallucination on Hallucination: Debate-Augmented RAG
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction
by: Jing, Xiaonan, et al.
Published: (2026)
by: Jing, Xiaonan, et al.
Published: (2026)
Similar Items
-
TravelBench : Exploring LLM Performance in Low-Resource Domains
by: Billa, Srinivas, et al.
Published: (2025) -
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025) -
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026) -
The Geometric Anatomy of Capability Acquisition in Transformers
by: Billa, Jayadev
Published: (2026) -
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026)