Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
Fuente:
arXiv
Salvato in:
| Autori principali: | Joo, Seongho, Min, Kyungmin, Koo, Jahyun, Jung, Kyomin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
di: Joo, Seongho, et al.
Pubblicazione: (2025)
di: Joo, Seongho, et al.
Pubblicazione: (2025)
Public Data Assisted Differentially Private In-Context Learning
di: Joo, Seongho, et al.
Pubblicazione: (2025)
di: Joo, Seongho, et al.
Pubblicazione: (2025)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
di: Xue, Yihao, et al.
Pubblicazione: (2025)
di: Xue, Yihao, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
di: Min, Kyungmin, et al.
Pubblicazione: (2024)
LLMs can be easily Confused by Instructional Distractions
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
di: Hwang, Yerin, et al.
Pubblicazione: (2025)
Program Synthesis via Test-Time Transduction
di: Lee, Kang-il, et al.
Pubblicazione: (2025)
di: Lee, Kang-il, et al.
Pubblicazione: (2025)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
Learned Hallucination Detection in Black-Box LLMs using Token-level Entropy Production Rate
di: Moslonka, Charles, et al.
Pubblicazione: (2025)
di: Moslonka, Charles, et al.
Pubblicazione: (2025)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
di: Kim, Junseok, et al.
Pubblicazione: (2026)
di: Kim, Junseok, et al.
Pubblicazione: (2026)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
di: Sawczyn, Albert, et al.
Pubblicazione: (2025)
Avoidance Decoding for Diverse Multi-Branch Story Generation
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
di: Park, Kyeongman, et al.
Pubblicazione: (2025)
LongStory: Coherent, Complete and Length Controlled Long story Generation
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
di: Park, Kyeongman, et al.
Pubblicazione: (2023)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
di: Kersting, Nicholas S., et al.
Pubblicazione: (2026)
di: Kersting, Nicholas S., et al.
Pubblicazione: (2026)
Enhancing Hallucination Detection via Future Context
di: Lee, Joosung, et al.
Pubblicazione: (2025)
di: Lee, Joosung, et al.
Pubblicazione: (2025)
Decoding Memories: An Efficient Pipeline for Self-Consistency Hallucination Detection
di: Gao, Weizhi, et al.
Pubblicazione: (2025)
di: Gao, Weizhi, et al.
Pubblicazione: (2025)
Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
SWITCH: Studying with Teacher for Knowledge Distillation of Large Language Models
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
di: Koo, Jahyun, et al.
Pubblicazione: (2024)
LifeTox: Unveiling Implicit Toxicity in Life Advice
di: Kim, Minbeom, et al.
Pubblicazione: (2023)
di: Kim, Minbeom, et al.
Pubblicazione: (2023)
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning
di: Jung, Min Jae, et al.
Pubblicazione: (2024)
di: Jung, Min Jae, et al.
Pubblicazione: (2024)
Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration
di: Mouzouni, Charafeddine
Pubblicazione: (2026)
di: Mouzouni, Charafeddine
Pubblicazione: (2026)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
Conditional [MASK] Discrete Diffusion Language Model
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
di: Koh, Hyukhun, et al.
Pubblicazione: (2024)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
When Wording Steers the Evaluation: Framing Bias in LLM judges
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
di: Hwang, Yerin, et al.
Pubblicazione: (2026)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
di: Yang, Nakyeong, et al.
Pubblicazione: (2023)
Hallucination Detection and Hallucination Mitigation: An Investigation
di: Luo, Junliang, et al.
Pubblicazione: (2024)
di: Luo, Junliang, et al.
Pubblicazione: (2024)
You've Changed: Detecting Modification of Black-Box Large Language Models
di: Dima, Alden, et al.
Pubblicazione: (2025)
di: Dima, Alden, et al.
Pubblicazione: (2025)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
di: Woo, Sangmin, et al.
Pubblicazione: (2025)
ReflectCAP: Detailed Image Captioning with Reflective Memory
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
Training Deliberative Monitors for Black-Box Scheming Detection
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
di: Sinha, Aditya, et al.
Pubblicazione: (2026)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
di: Hu, Mengya, et al.
Pubblicazione: (2024)
di: Hu, Mengya, et al.
Pubblicazione: (2024)
Cleanse: Uncertainty Estimation Approach Using Clustering-based Semantic Consistency in LLMs
di: Joo, Minsuh, et al.
Pubblicazione: (2025)
di: Joo, Minsuh, et al.
Pubblicazione: (2025)
HARP: Hallucination Detection via Reasoning Subspace Projection
di: Hu, Junjie, et al.
Pubblicazione: (2025)
di: Hu, Junjie, et al.
Pubblicazione: (2025)
Collaborative Stance Detection via Small-Large Language Model Consistency Verification
di: Yan, Yu, et al.
Pubblicazione: (2025)
di: Yan, Yu, et al.
Pubblicazione: (2025)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
di: Chakraborty, Neeloy, et al.
Pubblicazione: (2025)
di: Chakraborty, Neeloy, et al.
Pubblicazione: (2025)
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs
di: Hwang, Yerin, et al.
Pubblicazione: (2024)
di: Hwang, Yerin, et al.
Pubblicazione: (2024)
The Energy of Falsehood: Detecting Hallucinations via Diffusion Model Likelihoods
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
di: Gautam, Arpit Singh, et al.
Pubblicazione: (2026)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
di: Kim, Minsung, et al.
Pubblicazione: (2025)
di: Kim, Minsung, et al.
Pubblicazione: (2025)
Steer LLM Latents for Hallucination Detection
di: Park, Seongheon, et al.
Pubblicazione: (2025)
di: Park, Seongheon, et al.
Pubblicazione: (2025)
Black-Box On-Policy Distillation of Large Language Models
di: Ye, Tianzhu, et al.
Pubblicazione: (2025)
di: Ye, Tianzhu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding
di: Joo, Seongho, et al.
Pubblicazione: (2025) -
Public Data Assisted Differentially Private In-Context Learning
di: Joo, Seongho, et al.
Pubblicazione: (2025) -
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
di: Xue, Yihao, et al.
Pubblicazione: (2025) -
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
di: Min, Kyungmin, et al.
Pubblicazione: (2024) -
LLMs can be easily Confused by Instructional Distractions
di: Hwang, Yerin, et al.
Pubblicazione: (2025)