Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
Fuente:
arXiv
Salvato in:
| Autori principali: | Mao, Nathan, Kaushik, Varun, Shivkumar, Shreya, Sharafoleslami, Parham, Zhu, Kevin, Dev, Sunishchal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
di: More, Abhishek, et al.
Pubblicazione: (2025)
di: More, Abhishek, et al.
Pubblicazione: (2025)
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
di: Sahay, Kenji, et al.
Pubblicazione: (2025)
di: Sahay, Kenji, et al.
Pubblicazione: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026)
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
On Early Detection of Hallucinations in Factual Question Answering
di: Snyder, Ben, et al.
Pubblicazione: (2023)
di: Snyder, Ben, et al.
Pubblicazione: (2023)
CA-BED: Conversation-Aware Bayesian Experimental Design
di: Arnould, Daniel, et al.
Pubblicazione: (2026)
di: Arnould, Daniel, et al.
Pubblicazione: (2026)
LLM Factoscope: Uncovering LLMs' Factual Discernment through Inner States Analysis
di: He, Jinwen, et al.
Pubblicazione: (2023)
di: He, Jinwen, et al.
Pubblicazione: (2023)
Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
di: Lee, Seongmin, et al.
Pubblicazione: (2024)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
HalluLens: LLM Hallucination Benchmark
di: Bang, Yejin, et al.
Pubblicazione: (2025)
di: Bang, Yejin, et al.
Pubblicazione: (2025)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
di: Agrawal, Shriyansh, et al.
Pubblicazione: (2025)
di: Agrawal, Shriyansh, et al.
Pubblicazione: (2025)
Generating Benchmarks for Factuality Evaluation of Language Models
di: Muhlgay, Dor, et al.
Pubblicazione: (2023)
di: Muhlgay, Dor, et al.
Pubblicazione: (2023)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
di: Yu, Lei, et al.
Pubblicazione: (2024)
di: Yu, Lei, et al.
Pubblicazione: (2024)
KG-FPQ: Evaluating Factuality Hallucination in LLMs with Knowledge Graph-based False Premise Questions
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
di: Zhu, Yanxu, et al.
Pubblicazione: (2024)
On Mitigating Code LLM Hallucinations with API Documentation
di: Jain, Nihal, et al.
Pubblicazione: (2024)
di: Jain, Nihal, et al.
Pubblicazione: (2024)
AutoRAG-LoRA: Hallucination-Triggered Knowledge Retuning via Lightweight Adapters
di: Dwivedi, Kaushik, et al.
Pubblicazione: (2025)
di: Dwivedi, Kaushik, et al.
Pubblicazione: (2025)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
di: Liu, Emmy, et al.
Pubblicazione: (2026)
di: Liu, Emmy, et al.
Pubblicazione: (2026)
Hallucination Detection with the Internal Layers of LLMs
di: Preiß, Martin
Pubblicazione: (2025)
di: Preiß, Martin
Pubblicazione: (2025)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
di: Su, Weihang, et al.
Pubblicazione: (2024)
di: Su, Weihang, et al.
Pubblicazione: (2024)
AI Hallucinations: A Misnomer Worth Clarifying
di: Maleki, Negar, et al.
Pubblicazione: (2024)
di: Maleki, Negar, et al.
Pubblicazione: (2024)
Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes
di: Myers, Quintin, et al.
Pubblicazione: (2025)
di: Myers, Quintin, et al.
Pubblicazione: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
di: Batra, Shourya, et al.
Pubblicazione: (2025)
di: Batra, Shourya, et al.
Pubblicazione: (2025)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
di: Zhou, Hanzhang, et al.
Pubblicazione: (2024)
di: Zhou, Hanzhang, et al.
Pubblicazione: (2024)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
di: Li, Junyi, et al.
Pubblicazione: (2025)
di: Li, Junyi, et al.
Pubblicazione: (2025)
From Confidence to Collapse in LLM Factual Robustness
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
di: Fastowski, Alina, et al.
Pubblicazione: (2025)
HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
di: Uluoglakci, Cem, et al.
Pubblicazione: (2024)
di: Uluoglakci, Cem, et al.
Pubblicazione: (2024)
Hallucination Benchmark in Medical Visual Question Answering
di: Wu, Jinge, et al.
Pubblicazione: (2024)
di: Wu, Jinge, et al.
Pubblicazione: (2024)
Permutation-Consensus Listwise Judging for Robust Factuality Evaluation
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
di: Huang, Tianyi, et al.
Pubblicazione: (2026)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
InFact: Informativeness Alignment for Improved LLM Factuality
di: Cohen, Roi, et al.
Pubblicazione: (2025)
di: Cohen, Roi, et al.
Pubblicazione: (2025)
Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
di: Atasoy, I. F., et al.
Pubblicazione: (2026)
Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Models
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
di: Rahman, Subhey Sadi, et al.
Pubblicazione: (2025)
One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
di: Wang, Pengbo, et al.
Pubblicazione: (2025)
di: Wang, Pengbo, et al.
Pubblicazione: (2025)
Internalizing LLM Reasoning via Discovery and Replay of Latent Actions
di: Shi, Zhenning, et al.
Pubblicazione: (2026)
di: Shi, Zhenning, et al.
Pubblicazione: (2026)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
di: Lee, DongGeon, et al.
Pubblicazione: (2025)
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
di: Cheng, Aileen, et al.
Pubblicazione: (2025)
di: Cheng, Aileen, et al.
Pubblicazione: (2025)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
Mitigating LLM Hallucinations via Conformal Abstention
di: Yadkori, Yasin Abbasi, et al.
Pubblicazione: (2024)
di: Yadkori, Yasin Abbasi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
di: More, Abhishek, et al.
Pubblicazione: (2025) -
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
di: Sahay, Kenji, et al.
Pubblicazione: (2025) -
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
di: Thomas, Rohan Subramanian, et al.
Pubblicazione: (2026) -
Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024) -
On Early Detection of Hallucinations in Factual Question Answering
di: Snyder, Ben, et al.
Pubblicazione: (2023)