Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Hazel, Lamb, Tom A., Bibi, Adel, Torr, Philip, Gal, Yarin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Universal In-Context Approximation By Prompting Fully Recurrent Models
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
von: Aichberger, Lukas, et al.
Veröffentlicht: (2025)
von: Aichberger, Lukas, et al.
Veröffentlicht: (2025)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2023)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2023)
Prompting a Pretrained Transformer Can Be a Universal Approximator
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
Fine-tuning can cripple your foundation model; preserving features may be the solution
von: Mukhoti, Jishnu, et al.
Veröffentlicht: (2023)
von: Mukhoti, Jishnu, et al.
Veröffentlicht: (2023)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
von: Lamb, Tom A., et al.
Veröffentlicht: (2026)
von: Lamb, Tom A., et al.
Veröffentlicht: (2026)
Towards Certification of Uncertainty Calibration under Adversarial Attacks
von: Emde, Cornelius, et al.
Veröffentlicht: (2024)
von: Emde, Cornelius, et al.
Veröffentlicht: (2024)
How Ambiguous Are the Rationales for Natural Language Reasoning? A Simple Approach to Handling Rationale Uncertainty
von: Kim, Hazel H.
Veröffentlicht: (2024)
von: Kim, Hazel H.
Veröffentlicht: (2024)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
von: Kossen, Jannik, et al.
Veröffentlicht: (2024)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2025)
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
von: Lin, Runqi, et al.
Veröffentlicht: (2025)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2024)
Safer by Diffusion, Broken by Context: Diffusion LLM's Safety Blessing and Its Failure Mode
von: He, Zeyuan, et al.
Veröffentlicht: (2026)
von: He, Zeyuan, et al.
Veröffentlicht: (2026)
Efficient Error Certification for Physics-Informed Neural Networks
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
von: Ouyang, Jialin
Veröffentlicht: (2025)
von: Ouyang, Jialin
Veröffentlicht: (2025)
Estimating the Hallucination Rate of Generative AI
von: Jesson, Andrew, et al.
Veröffentlicht: (2024)
von: Jesson, Andrew, et al.
Veröffentlicht: (2024)
Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
von: Yang, Yibo, et al.
Veröffentlicht: (2024)
von: Yang, Yibo, et al.
Veröffentlicht: (2024)
From Categories to Classifiers: Name-Only Continual Learning by Exploring the Web
von: Prabhu, Ameya, et al.
Veröffentlicht: (2023)
von: Prabhu, Ameya, et al.
Veröffentlicht: (2023)
Simple Baselines are Competitive with Code Evolution
von: Gideoni, Yonatan, et al.
Veröffentlicht: (2026)
von: Gideoni, Yonatan, et al.
Veröffentlicht: (2026)
The Benefits and Risks of Transductive Approaches for AI Fairness
von: Razzak, Muhammed, et al.
Veröffentlicht: (2024)
von: Razzak, Muhammed, et al.
Veröffentlicht: (2024)
OMNI-LEAK: Orchestrator Multi-Agent Network Induced Data Leakage
von: Naik, Akshat, et al.
Veröffentlicht: (2026)
von: Naik, Akshat, et al.
Veröffentlicht: (2026)
Focus On This, Not That! Steering LLMs with Adaptive Feature Specification
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
von: Lamb, Tom A., et al.
Veröffentlicht: (2024)
Do Multilingual LLMs Think In English?
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
Mixture of Experts Made Intrinsically Interpretable
von: Yang, Xingyi, et al.
Veröffentlicht: (2025)
von: Yang, Xingyi, et al.
Veröffentlicht: (2025)
Temporal-Difference Variational Continual Learning
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
von: Melo, Luckeciano C., et al.
Veröffentlicht: (2024)
Scaling Up Active Testing to Large Language Models
von: Berrada, Gabrielle, et al.
Veröffentlicht: (2025)
von: Berrada, Gabrielle, et al.
Veröffentlicht: (2025)
Shh, don't say that! Domain Certification in LLMs
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
von: Emde, Cornelius, et al.
Veröffentlicht: (2025)
On Pretraining Data Diversity for Self-Supervised Learning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
von: Sneh, Jonathan, et al.
Veröffentlicht: (2025)
LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Training
von: Gwak, Minju, et al.
Veröffentlicht: (2026)
von: Gwak, Minju, et al.
Veröffentlicht: (2026)
TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
von: Chang, Shenxu, et al.
Veröffentlicht: (2025)
von: Chang, Shenxu, et al.
Veröffentlicht: (2025)
Richer Bayesian Last Layers with Subsampled NTK Features
von: Calvo-Ordoñez, Sergio, et al.
Veröffentlicht: (2026)
von: Calvo-Ordoñez, Sergio, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Universal In-Context Approximation By Prompting Fully Recurrent Models
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024) -
MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents
von: Aichberger, Lukas, et al.
Veröffentlicht: (2025) -
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2023) -
Prompting a Pretrained Transformer Can Be a Universal Approximator
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024) -
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)