Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kossen, Jannik, Han, Jiatong, Razzak, Muhammed, Schut, Lisa, Malik, Shreshth, Gal, Yarin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
di: Nikitin, Alexander, et al.
Pubblicazione: (2024)
di: Nikitin, Alexander, et al.
Pubblicazione: (2024)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
di: Tjandra, Benedict Aaron, et al.
Pubblicazione: (2024)
di: Tjandra, Benedict Aaron, et al.
Pubblicazione: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
di: Kossen, Jannik, et al.
Pubblicazione: (2023)
Do Multilingual LLMs Think In English?
di: Schut, Lisa, et al.
Pubblicazione: (2025)
di: Schut, Lisa, et al.
Pubblicazione: (2025)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
di: Penny-Dimri, Jahan C., et al.
Pubblicazione: (2025)
di: Penny-Dimri, Jahan C., et al.
Pubblicazione: (2025)
Scaling Up Active Testing to Large Language Models
di: Berrada, Gabrielle, et al.
Pubblicazione: (2025)
di: Berrada, Gabrielle, et al.
Pubblicazione: (2025)
Iterative Deployment Improves Planning Skills in LLMs
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors
di: Nicolson, Angus, et al.
Pubblicazione: (2024)
di: Nicolson, Angus, et al.
Pubblicazione: (2024)
Cost-Effective Hallucination Detection for LLMs
di: Valentin, Simon, et al.
Pubblicazione: (2024)
di: Valentin, Simon, et al.
Pubblicazione: (2024)
TABES: Trajectory-Aware Backward-on-Entropy Steering for Masked Diffusion Models
di: Saini, Shreshth, et al.
Pubblicazione: (2026)
di: Saini, Shreshth, et al.
Pubblicazione: (2026)
The Benefits and Risks of Transductive Approaches for AI Fairness
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
di: Razzak, Muhammed, et al.
Pubblicazione: (2024)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
di: Janiak, Denis, et al.
Pubblicazione: (2025)
di: Janiak, Denis, et al.
Pubblicazione: (2025)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
di: Frank, Gregory N.
Pubblicazione: (2026)
di: Frank, Gregory N.
Pubblicazione: (2026)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
di: Abdulhai, Marwa, et al.
Pubblicazione: (2025)
di: Abdulhai, Marwa, et al.
Pubblicazione: (2025)
Towards a Neural Debugger for Python
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
di: Li, Qing, et al.
Pubblicazione: (2025)
di: Li, Qing, et al.
Pubblicazione: (2025)
Hallucination Detection in LLMs Using Spectral Features of Attention Maps
di: Binkowski, Jakub, et al.
Pubblicazione: (2025)
di: Binkowski, Jakub, et al.
Pubblicazione: (2025)
MADE: Benchmark Environments for Closed-Loop Materials Discovery
di: Malik, Shreshth A, et al.
Pubblicazione: (2026)
di: Malik, Shreshth A, et al.
Pubblicazione: (2026)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
Hallucination Detection in LLMs: Fast and Memory-Efficient Fine-Tuned Models
di: Arteaga, Gabriel Y., et al.
Pubblicazione: (2024)
di: Arteaga, Gabriel Y., et al.
Pubblicazione: (2024)
The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
di: Sinha, Debu
Pubblicazione: (2025)
di: Sinha, Debu
Pubblicazione: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
di: Wang, Shaowen, et al.
Pubblicazione: (2025)
Estimating the Hallucination Rate of Generative AI
di: Jesson, Andrew, et al.
Pubblicazione: (2024)
di: Jesson, Andrew, et al.
Pubblicazione: (2024)
Simple Baselines are Competitive with Code Evolution
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
di: Gideoni, Yonatan, et al.
Pubblicazione: (2026)
Learning to Reason for Hallucination Span Detection
di: Su, Hsuan, et al.
Pubblicazione: (2025)
di: Su, Hsuan, et al.
Pubblicazione: (2025)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
di: Melo, Luckeciano C., et al.
Pubblicazione: (2025)
Temporal-Difference Variational Continual Learning
di: Melo, Luckeciano C., et al.
Pubblicazione: (2024)
di: Melo, Luckeciano C., et al.
Pubblicazione: (2024)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
di: Cho, Nicole, et al.
Pubblicazione: (2025)
di: Cho, Nicole, et al.
Pubblicazione: (2025)
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2025)
Semantic Faithfulness and Entropy Production Measures to Tame Your LLM Demons and Manage Hallucinations
di: Halperin, Igor
Pubblicazione: (2025)
di: Halperin, Igor
Pubblicazione: (2025)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
di: Le, Qi, et al.
Pubblicazione: (2025)
di: Le, Qi, et al.
Pubblicazione: (2025)
Hallucinated Span Detection with Multi-View Attention Features
di: Ogasa, Yuya, et al.
Pubblicazione: (2025)
di: Ogasa, Yuya, et al.
Pubblicazione: (2025)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
di: Gao, Cheng, et al.
Pubblicazione: (2026)
di: Gao, Cheng, et al.
Pubblicazione: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
di: Olson, Matthew Lyle, et al.
Pubblicazione: (2025)
di: Olson, Matthew Lyle, et al.
Pubblicazione: (2025)
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
di: Kowshik, Suhas S, et al.
Pubblicazione: (2024)
di: Kowshik, Suhas S, et al.
Pubblicazione: (2024)
AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration - Learning from Cheap, Optimizing Expensive
di: Guo, Taicheng, et al.
Pubblicazione: (2026)
di: Guo, Taicheng, et al.
Pubblicazione: (2026)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
di: Halperin, Igor
Pubblicazione: (2025)
di: Halperin, Igor
Pubblicazione: (2025)
Documenti analoghi
-
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
di: Nikitin, Alexander, et al.
Pubblicazione: (2024) -
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
di: Tjandra, Benedict Aaron, et al.
Pubblicazione: (2024) -
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
di: Kossen, Jannik, et al.
Pubblicazione: (2023) -
Do Multilingual LLMs Think In English?
di: Schut, Lisa, et al.
Pubblicazione: (2025) -
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
di: Penny-Dimri, Jahan C., et al.
Pubblicazione: (2025)