Distinguishing Ignorance from Error in LLM Hallucinations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Simhi, Adi, Herzig, Jonathan, Szpektor, Idan, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
par: Simhi, Adi, et autres
Publié: (2024)
par: Simhi, Adi, et autres
Publié: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
HACK: Hallucinations Along Certainty and Knowledge Axes
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
par: Simhi, Adi, et autres
Publié: (2025)
par: Simhi, Adi, et autres
Publié: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
par: Orgad, Hadas, et autres
Publié: (2024)
par: Orgad, Hadas, et autres
Publié: (2024)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
par: Simhi, Adi, et autres
Publié: (2026)
par: Simhi, Adi, et autres
Publié: (2026)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
par: Ashuach, Tomer, et autres
Publié: (2025)
par: Ashuach, Tomer, et autres
Publié: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
par: Ashuach, Tomer, et autres
Publié: (2024)
par: Ashuach, Tomer, et autres
Publié: (2024)
Are formal and functional linguistic mechanisms dissociated in language models?
par: Hanna, Michael, et autres
Publié: (2025)
par: Hanna, Michael, et autres
Publié: (2025)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
par: Ashuach, Tomer, et autres
Publié: (2026)
par: Ashuach, Tomer, et autres
Publié: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
par: Hanna, Michael, et autres
Publié: (2024)
par: Hanna, Michael, et autres
Publié: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
par: Toker, Michael, et autres
Publié: (2024)
par: Toker, Michael, et autres
Publié: (2024)
Reasoning Models Know What's Important, and Encode It in Their Activations
par: Nikankin, Yaniv, et autres
Publié: (2026)
par: Nikankin, Yaniv, et autres
Publié: (2026)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
par: Nikankin, Yaniv, et autres
Publié: (2024)
par: Nikankin, Yaniv, et autres
Publié: (2024)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
par: Nikankin, Yaniv, et autres
Publié: (2025)
par: Nikankin, Yaniv, et autres
Publié: (2025)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
par: Arad, Dana, et autres
Publié: (2023)
par: Arad, Dana, et autres
Publié: (2023)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
par: Rudman, William, et autres
Publié: (2026)
par: Rudman, William, et autres
Publié: (2026)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
par: Toker, Michael, et autres
Publié: (2024)
par: Toker, Michael, et autres
Publié: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
par: Oketunji, Abiodun Finbarrs
Publié: (2023)
Position-aware Automatic Circuit Discovery
par: Haklay, Tal, et autres
Publié: (2025)
par: Haklay, Tal, et autres
Publié: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
par: Zhu, Qian, et autres
Publié: (2026)
par: Zhu, Qian, et autres
Publié: (2026)
Extracting Structured Insights from Financial News: An Augmented LLM Driven Approach
par: Dolphin, Rian, et autres
Publié: (2024)
par: Dolphin, Rian, et autres
Publié: (2024)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
par: Rathva, Harsh, et autres
Publié: (2025)
par: Rathva, Harsh, et autres
Publié: (2025)
A Benchmark of French ASR Systems Based on Error Severity
par: Tholly, Antoine, et autres
Publié: (2025)
par: Tholly, Antoine, et autres
Publié: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
par: Orgad, Hadas, et autres
Publié: (2026)
par: Orgad, Hadas, et autres
Publié: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
par: Liu, Zhongxin, et autres
Publié: (2025)
par: Liu, Zhongxin, et autres
Publié: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
par: Cherif, Ahmed
Publié: (2026)
par: Cherif, Ahmed
Publié: (2026)
Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs
par: Ovcharov, Volodymyr
Publié: (2026)
par: Ovcharov, Volodymyr
Publié: (2026)
VertAttack: Taking advantage of Text Classifiers' horizontal vision
par: Rusert, Jonathan
Publié: (2024)
par: Rusert, Jonathan
Publié: (2024)
RedHerring Attack: Testing the Reliability of Attack Detection
par: Rusert, Jonathan
Publié: (2025)
par: Rusert, Jonathan
Publié: (2025)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
par: Karinshak, Elise, et autres
Publié: (2024)
par: Karinshak, Elise, et autres
Publié: (2024)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
par: Wang, Xintao, et autres
Publié: (2026)
par: Wang, Xintao, et autres
Publié: (2026)
A Multi-Pass Large Language Model Framework for Precise and Efficient Radiology Report Error Detection
par: Kim, Songsoo, et autres
Publié: (2025)
par: Kim, Songsoo, et autres
Publié: (2025)
Revisiting Word Embeddings in the LLM Era
par: Mahajan, Yash, et autres
Publié: (2024)
par: Mahajan, Yash, et autres
Publié: (2024)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
par: Williamson, Dane, et autres
Publié: (2025)
par: Williamson, Dane, et autres
Publié: (2025)
StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis
par: Opara, Chidimma
Publié: (2024)
par: Opara, Chidimma
Publié: (2024)
Decoding-Free Sampling Strategies for LLM Marginalization
par: Pohl, David, et autres
Publié: (2025)
par: Pohl, David, et autres
Publié: (2025)
LUCID: LLM-Generated Utterances for Complex and Interesting Dialogues
par: Stacey, Joe, et autres
Publié: (2024)
par: Stacey, Joe, et autres
Publié: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
par: Ge, Danying, et autres
Publié: (2025)
par: Ge, Danying, et autres
Publié: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
par: Tu, Songjun, et autres
Publié: (2026)
par: Tu, Songjun, et autres
Publié: (2026)
Documents similaires
-
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
par: Simhi, Adi, et autres
Publié: (2024) -
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
par: Simhi, Adi, et autres
Publié: (2025) -
HACK: Hallucinations Along Certainty and Knowledge Axes
par: Simhi, Adi, et autres
Publié: (2025) -
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
par: Simhi, Adi, et autres
Publié: (2025) -
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
par: Orgad, Hadas, et autres
Publié: (2024)