Can LLMs Detect Their Own Hallucinations?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kadotani, Sora, Nishida, Kosuke, Nishida, Kyosuke |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
par: Nishida, Kosuke, et autres
Publié: (2024)
par: Nishida, Kosuke, et autres
Publié: (2024)
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
par: Shinoda, Kazutoshi, et autres
Publié: (2026)
par: Shinoda, Kazutoshi, et autres
Publié: (2026)
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
par: Kawasaki, Haruka, et autres
Publié: (2026)
par: Kawasaki, Haruka, et autres
Publié: (2026)
Wavelet-based Positional Representation for Long Context
par: Oka, Yui, et autres
Publié: (2025)
par: Oka, Yui, et autres
Publié: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
par: Tanaka, Ryota, et autres
Publié: (2024)
par: Tanaka, Ryota, et autres
Publié: (2024)
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
par: Shinoda, Kazutoshi, et autres
Publié: (2025)
par: Shinoda, Kazutoshi, et autres
Publié: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
par: Tanaka, Ryota, et autres
Publié: (2025)
par: Tanaka, Ryota, et autres
Publié: (2025)
LLMs Can Generate a Better Answer by Aggregating Their Own Responses
par: Li, Zichong, et autres
Publié: (2025)
par: Li, Zichong, et autres
Publié: (2025)
Let's Put Ourselves in Sally's Shoes: Shoes-of-Others Prefilling Improves Theory of Mind in Large Language Models
par: Shinoda, Kazutoshi, et autres
Publié: (2025)
par: Shinoda, Kazutoshi, et autres
Publié: (2025)
Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
par: Gogoulou, Evangelia, et autres
Publié: (2025)
par: Gogoulou, Evangelia, et autres
Publié: (2025)
Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
par: Ghasemabadi, Amirhosein, et autres
Publié: (2025)
par: Ghasemabadi, Amirhosein, et autres
Publié: (2025)
Instability in Downstream Task Performance During LLM Pretraining
par: Nishida, Yuto, et autres
Publié: (2025)
par: Nishida, Yuto, et autres
Publié: (2025)
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
par: Kamoi, Ryo, et autres
Publié: (2024)
par: Kamoi, Ryo, et autres
Publié: (2024)
Critical Confabulation: Can LLMs Hallucinate for Social Good?
par: Sui, Peiqi, et autres
Publié: (2025)
par: Sui, Peiqi, et autres
Publié: (2025)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
par: Chijiwa, Daiki, et autres
Publié: (2025)
par: Chijiwa, Daiki, et autres
Publié: (2025)
Can Hallucinations Help? Boosting LLMs for Drug Discovery
par: Yuan, Shuzhou, et autres
Publié: (2025)
par: Yuan, Shuzhou, et autres
Publié: (2025)
Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
par: Phillips, Edward, et autres
Publié: (2025)
par: Phillips, Edward, et autres
Publié: (2025)
Generating Diverse Translation with Perturbed kNN-MT
par: Nishida, Yuto, et autres
Publié: (2024)
par: Nishida, Yuto, et autres
Publié: (2024)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
par: Nishida, Naoto, et autres
Publié: (2025)
par: Nishida, Naoto, et autres
Publié: (2025)
BYOL: Bring Your Own Language Into LLMs
par: Zamir, Syed Waqas, et autres
Publié: (2026)
par: Zamir, Syed Waqas, et autres
Publié: (2026)
Hallucination Detection with the Internal Layers of LLMs
par: Preiß, Martin
Publié: (2025)
par: Preiß, Martin
Publié: (2025)
Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
par: Feroz, Faisal, et autres
Publié: (2026)
par: Feroz, Faisal, et autres
Publié: (2026)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
par: Agrawal, Garima, et autres
Publié: (2023)
par: Agrawal, Garima, et autres
Publié: (2023)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
par: Qi, Siya, et autres
Publié: (2026)
par: Qi, Siya, et autres
Publié: (2026)
How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments
par: Ide, Yusuke, et autres
Publié: (2024)
par: Ide, Yusuke, et autres
Publié: (2024)
Post Persona Alignment for Multi-Session Dialogue Generation
par: Chen, Yi-Pei, et autres
Publié: (2025)
par: Chen, Yi-Pei, et autres
Publié: (2025)
Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
par: Takahashi, Hiroshi, et autres
Publié: (2026)
par: Takahashi, Hiroshi, et autres
Publié: (2026)
Explanation Bottleneck Models
par: Yamaguchi, Shin'ya, et autres
Publié: (2024)
par: Yamaguchi, Shin'ya, et autres
Publié: (2024)
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
par: Peisakhovsky, Yehonatan, et autres
Publié: (2025)
par: Peisakhovsky, Yehonatan, et autres
Publié: (2025)
INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
par: Chen, Chao, et autres
Publié: (2024)
par: Chen, Chao, et autres
Publié: (2024)
Quantifying and Mitigating Socially Desirable Responding in LLMs: A Desirability-Matched Graded Forced-Choice Psychometric Study
par: Okada, Kensuke, et autres
Publié: (2026)
par: Okada, Kensuke, et autres
Publié: (2026)
Cost-Effective Hallucination Detection for LLMs
par: Valentin, Simon, et autres
Publié: (2024)
par: Valentin, Simon, et autres
Publié: (2024)
Do LLMs Benefit From Their Own Words?
par: Huang, Jenny Y., et autres
Publié: (2026)
par: Huang, Jenny Y., et autres
Publié: (2026)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
par: Bui, Anh Thu Maria, et autres
Publié: (2024)
par: Bui, Anh Thu Maria, et autres
Publié: (2024)
Long-Tail Crisis in Nearest Neighbor Language Models
par: Nishida, Yuto, et autres
Publié: (2025)
par: Nishida, Yuto, et autres
Publié: (2025)
DiffNator: Generating Structured Explanations of Time-Series Differences
par: Dohi, Kota, et autres
Publié: (2025)
par: Dohi, Kota, et autres
Publié: (2025)
Retrieving Time-Series Differences Using Natural Language Queries
par: Dohi, Kota, et autres
Publié: (2025)
par: Dohi, Kota, et autres
Publié: (2025)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
SINdex: Semantic INconsistency Index for Hallucination Detection in LLMs
par: Abdaljalil, Samir, et autres
Publié: (2025)
par: Abdaljalil, Samir, et autres
Publié: (2025)
Hallucination Detection in LLMs with Topological Divergence on Attention Graphs
par: Bazarova, Alexandra, et autres
Publié: (2025)
par: Bazarova, Alexandra, et autres
Publié: (2025)
Documents similaires
-
Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes
par: Nishida, Kosuke, et autres
Publié: (2024) -
Debiasing Reward Models via Causally Motivated Inference-Time Intervention
par: Shinoda, Kazutoshi, et autres
Publié: (2026) -
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
par: Kawasaki, Haruka, et autres
Publié: (2026) -
Wavelet-based Positional Representation for Long Context
par: Oka, Yui, et autres
Publié: (2025) -
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
par: Tanaka, Ryota, et autres
Publié: (2024)