Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Yihao, Greenewald, Kristjan, Mroueh, Youssef, Mirzasoleiman, Baharan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
von: Joo, Seongho, et al.
Veröffentlicht: (2025)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
von: Gao, Weizhi, et al.
Veröffentlicht: (2025)
von: Gao, Weizhi, et al.
Veröffentlicht: (2025)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
von: Moslonka, Charles, et al.
Veröffentlicht: (2025)
von: Moslonka, Charles, et al.
Veröffentlicht: (2025)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2024)
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2024)
Large Language Models can be Strong Self-Detoxifiers
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
Hallucination Detection and Hallucination Mitigation: An Investigation
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
von: Luo, Junliang, et al.
Veröffentlicht: (2024)
You've Changed: Detecting Modification of Black-Box Large Language Models
von: Dima, Alden, et al.
Veröffentlicht: (2025)
von: Dima, Alden, et al.
Veröffentlicht: (2025)
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
von: Chen, Yuefei, et al.
Veröffentlicht: (2026)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Training Deliberative Monitors for Black-Box Scheming Detection
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
von: Hu, Mengya, et al.
Veröffentlicht: (2024)
von: Hu, Mengya, et al.
Veröffentlicht: (2024)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
von: Brüel-Gabrielsson, Rickard, et al.
Veröffentlicht: (2024)
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language Models
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
von: Liu, Qiang, et al.
Veröffentlicht: (2025)
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
von: Li, Zhuo, et al.
Veröffentlicht: (2024)
Domain Adaptable Prescriptive AI Agent for Enterprise
von: Orderique, Piero, et al.
Veröffentlicht: (2024)
von: Orderique, Piero, et al.
Veröffentlicht: (2024)
Steering the Verifiability of Multimodal AI Hallucinations
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
von: Pang, Jianhong, et al.
Veröffentlicht: (2026)
Black-Box On-Policy Distillation of Large Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
A Survey of Calibration Process for Black-Box LLMs
von: Xie, Liangru, et al.
Veröffentlicht: (2024)
von: Xie, Liangru, et al.
Veröffentlicht: (2024)
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks
von: Rair, Nisrine, et al.
Veröffentlicht: (2026)
von: Rair, Nisrine, et al.
Veröffentlicht: (2026)
Beyond Facts: Evaluating Intent Hallucination in Large Language Models
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
von: Hao, Yijie, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice
von: Doshi, Savan
Veröffentlicht: (2026)
von: Doshi, Savan
Veröffentlicht: (2026)
Hallucination Detection with Small Language Models
von: Cheung, Ming
Veröffentlicht: (2025)
von: Cheung, Ming
Veröffentlicht: (2025)
Hallucination Detection with the Internal Layers of LLMs
von: Preiß, Martin
Veröffentlicht: (2025)
von: Preiß, Martin
Veröffentlicht: (2025)
Ähnliche Einträge
-
Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
von: Joo, Seongho, et al.
Veröffentlicht: (2025) -
LoRA is All You Need for Safety Alignment of Reasoning LLMs
von: Xue, Yihao, et al.
Veröffentlicht: (2025) -
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
von: Xue, Yihao, et al.
Veröffentlicht: (2025) -
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
von: Yang, Yu, et al.
Veröffentlicht: (2024) -
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)