Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Simhi, Adi, Herzig, Jonathan, Szpektor, Idan, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
HACK: Hallucinations Along Certainty and Knowledge Axes
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
von: Ashuach, Tomer, et al.
Veröffentlicht: (2024)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2024)
Are formal and functional linguistic mechanisms dissociated in language models?
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
von: Hanna, Michael, et al.
Veröffentlicht: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
von: Ashuach, Tomer, et al.
Veröffentlicht: (2026)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2026)
Reasoning Models Know What's Important, and Encode It in Their Activations
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2026)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2026)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
von: Cherif, Ahmed
Veröffentlicht: (2026)
von: Cherif, Ahmed
Veröffentlicht: (2026)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
von: Arad, Dana, et al.
Veröffentlicht: (2023)
von: Arad, Dana, et al.
Veröffentlicht: (2023)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
von: Rudman, William, et al.
Veröffentlicht: (2026)
von: Rudman, William, et al.
Veröffentlicht: (2026)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Position-aware Automatic Circuit Discovery
von: Haklay, Tal, et al.
Veröffentlicht: (2025)
von: Haklay, Tal, et al.
Veröffentlicht: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
VertAttack: Taking advantage of Text Classifiers' horizontal vision
von: Rusert, Jonathan
Veröffentlicht: (2024)
von: Rusert, Jonathan
Veröffentlicht: (2024)
RedHerring Attack: Testing the Reliability of Attack Detection
von: Rusert, Jonathan
Veröffentlicht: (2025)
von: Rusert, Jonathan
Veröffentlicht: (2025)
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2024)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
PL-Guard: Benchmarking Language Model Safety for Polish
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
von: Qi, Jinhu, et al.
Veröffentlicht: (2024)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
von: Cui, Hyang
Veröffentlicht: (2025)
von: Cui, Hyang
Veröffentlicht: (2025)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
von: Chehbouni, Khaoula, et al.
Veröffentlicht: (2025)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
von: Zheng, Qinyue, et al.
Veröffentlicht: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
A Benchmark of French ASR Systems Based on Error Severity
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
LLMs and the Human Condition
von: Wallis, Peter
Veröffentlicht: (2024)
von: Wallis, Peter
Veröffentlicht: (2024)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024) -
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025) -
HACK: Hallucinations Along Certainty and Knowledge Axes
von: Simhi, Adi, et al.
Veröffentlicht: (2025) -
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025) -
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)