Estimating the Hallucination Rate of Generative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Jesson, Andrew, Beltran-Velez, Nicolas, Chu, Quentin, Karlekar, Sweta, Kossen, Jannik, Gal, Yarin, Cunningham, John P., Blei, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023)
by: Kossen, Jannik, et al.
Published: (2023)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024)
by: Nikitin, Alexander, et al.
Published: (2024)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024)
by: Kossen, Jannik, et al.
Published: (2024)
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)
by: Karlekar, Sweta, et al.
Published: (2026)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
by: Tjandra, Benedict Aaron, et al.
Published: (2024)
ReLU to the Rescue: Improve Your On-Policy Actor-Critic with Positive Advantages
by: Jesson, Andrew, et al.
Published: (2023)
by: Jesson, Andrew, et al.
Published: (2023)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
by: Zheng, Carolina, et al.
Published: (2025)
by: Zheng, Carolina, et al.
Published: (2025)
Scaling Up Active Testing to Large Language Models
by: Berrada, Gabrielle, et al.
Published: (2025)
by: Berrada, Gabrielle, et al.
Published: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
by: Shi, Claudia, et al.
Published: (2024)
by: Shi, Claudia, et al.
Published: (2024)
Improving Generalization on the ProcGen Benchmark with Simple Architectural Changes and Scale
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
Treeffuser: Probabilistic Predictions via Conditional Diffusions with Gradient-Boosted Trees
by: Beltran-Velez, Nicolas, et al.
Published: (2024)
by: Beltran-Velez, Nicolas, et al.
Published: (2024)
The Benefits and Risks of Transductive Approaches for AI Fairness
by: Razzak, Muhammed, et al.
Published: (2024)
by: Razzak, Muhammed, et al.
Published: (2024)
Towards a Neural Debugger for Python
by: Beck, Maximilian, et al.
Published: (2026)
by: Beck, Maximilian, et al.
Published: (2026)
Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions
by: Kim, Hazel, et al.
Published: (2024)
by: Kim, Hazel, et al.
Published: (2024)
Bayesian Invariance Modeling of Multi-Environment Data
by: Wu, Luhuan, et al.
Published: (2025)
by: Wu, Luhuan, et al.
Published: (2025)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
by: Penny-Dimri, Jahan C., et al.
Published: (2025)
by: Penny-Dimri, Jahan C., et al.
Published: (2025)
Density Uncertainty Layers for Reliable Uncertainty Estimation
by: Park, Yookoon, et al.
Published: (2023)
by: Park, Yookoon, et al.
Published: (2023)
Simple Baselines are Competitive with Code Evolution
by: Gideoni, Yonatan, et al.
Published: (2026)
by: Gideoni, Yonatan, et al.
Published: (2026)
Practical and Asymptotically Exact Conditional Sampling in Diffusion Models
by: Wu, Luhuan, et al.
Published: (2023)
by: Wu, Luhuan, et al.
Published: (2023)
Optimization-based Causal Estimation from Heterogenous Environments
by: Yin, Mingzhang, et al.
Published: (2021)
by: Yin, Mingzhang, et al.
Published: (2021)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Temporal-Difference Variational Continual Learning
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
JULI: Jailbreak Large Language Models by Self-Introspection
by: Wang, Jesson, et al.
Published: (2025)
by: Wang, Jesson, et al.
Published: (2025)
Rethinking Aleatoric and Epistemic Uncertainty
by: Smith, Freddie Bickford, et al.
Published: (2024)
by: Smith, Freddie Bickford, et al.
Published: (2024)
Extremely Greedy Equivalence Search
by: Nazaret, Achille, et al.
Published: (2025)
by: Nazaret, Achille, et al.
Published: (2025)
Estimating Wage Disparities Using Foundation Models
by: Vafa, Keyon, et al.
Published: (2024)
by: Vafa, Keyon, et al.
Published: (2024)
Estimating the Causal Effects of T Cell Receptors
by: Weinstein, Eli N., et al.
Published: (2024)
by: Weinstein, Eli N., et al.
Published: (2024)
Robust Representation Learning through Explicit Environment Modeling
by: Slavutsky, Yuli, et al.
Published: (2026)
by: Slavutsky, Yuli, et al.
Published: (2026)
Quantifying Uncertainty in the Presence of Distribution Shifts
by: Slavutsky, Yuli, et al.
Published: (2025)
by: Slavutsky, Yuli, et al.
Published: (2025)
Extending Mean-Field Variational Inference via Entropic Regularization: Theory and Computation
by: Wu, Bohan, et al.
Published: (2024)
by: Wu, Bohan, et al.
Published: (2024)
Neural Generalized Mixed-Effects Models
by: Slavutsky, Yuli, et al.
Published: (2026)
by: Slavutsky, Yuli, et al.
Published: (2026)
Training Transformers for KV Cache Compressibility
by: Gelberg, Yoav, et al.
Published: (2026)
by: Gelberg, Yoav, et al.
Published: (2026)
Deep Bayesian Active Learning for Preference Modeling in Large Language Models
by: Melo, Luckeciano C., et al.
Published: (2024)
by: Melo, Luckeciano C., et al.
Published: (2024)
Amortized Variational Inference: When and Why?
by: Margossian, Charles C., et al.
Published: (2023)
by: Margossian, Charles C., et al.
Published: (2023)
The Curse of Recursion: Training on Generated Data Makes Models Forget
by: Shumailov, Ilia, et al.
Published: (2023)
by: Shumailov, Ilia, et al.
Published: (2023)
Leveraging Deep Learning for Physical Model Bias of Global Air Quality Estimates
by: Doerksen, Kelsey, et al.
Published: (2025)
by: Doerksen, Kelsey, et al.
Published: (2025)
Hierarchical Causal Models
by: Weinstein, Eli N., et al.
Published: (2024)
by: Weinstein, Eli N., et al.
Published: (2024)
Boundary Point Jailbreaking of Black-Box LLMs
by: Davies, Xander, et al.
Published: (2026)
by: Davies, Xander, et al.
Published: (2026)
Similar Items
-
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024) -
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
by: Kossen, Jannik, et al.
Published: (2023) -
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
by: Nikitin, Alexander, et al.
Published: (2024) -
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
by: Kossen, Jannik, et al.
Published: (2024) -
Duel-Evolve: Reward-Free Test-Time Scaling via LLM Self-Preferences
by: Karlekar, Sweta, et al.
Published: (2026)