Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
Fuente:
arXiv
Saved in:
| Main Authors: | Lamba, Naveen, Tiwari, Sanju, Gaur, Manas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025)
by: Lamba, Naveen, et al.
Published: (2025)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024)
by: Zhu, Zhiying, et al.
Published: (2024)
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024)
by: Wang, Binjie, et al.
Published: (2024)
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
by: Chen, Kedi, et al.
Published: (2024)
by: Chen, Kedi, et al.
Published: (2024)
KnowHalu: Hallucination Detection via Multi-Form Knowledge Based Factual Checking
by: Zhang, Jiawei, et al.
Published: (2024)
by: Zhang, Jiawei, et al.
Published: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
by: Mohammadi, Seyedali, et al.
Published: (2026)
by: Mohammadi, Seyedali, et al.
Published: (2026)
CodeGemma: Open Code Models Based on Gemma
by: CodeGemma Team, et al.
Published: (2024)
by: CodeGemma Team, et al.
Published: (2024)
Entropy and Attention Dynamics in Small Language Models: A Trace-Level Structural Analysis on the TruthfulQA Benchmark
by: Adeseye, Adeyemi, et al.
Published: (2026)
by: Adeseye, Adeyemi, et al.
Published: (2026)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
by: Zhang, Shaolei, et al.
Published: (2024)
by: Zhang, Shaolei, et al.
Published: (2024)
Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
by: Luo, Wen, et al.
Published: (2026)
by: Luo, Wen, et al.
Published: (2026)
Neurosymbolic Retrievers for Retrieval-augmented Generation
by: Saxena, Yash, et al.
Published: (2026)
by: Saxena, Yash, et al.
Published: (2026)
Triggering Hallucinations in LLMs: A Quantitative Study of Prompt-Induced Hallucination in Large Language Models
by: Sato, Makoto
Published: (2025)
by: Sato, Makoto
Published: (2025)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
by: Agrawal, Tanmay
Published: (2025)
by: Agrawal, Tanmay
Published: (2025)
Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
Hallucination Detection and Hallucination Mitigation: An Investigation
by: Luo, Junliang, et al.
Published: (2024)
by: Luo, Junliang, et al.
Published: (2024)
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024)
by: Mohammadi, Seyedali, et al.
Published: (2024)
Gemma 3 Technical Report
by: Gemma Team, et al.
Published: (2025)
by: Gemma Team, et al.
Published: (2025)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Gemma: Open Models Based on Gemini Research and Technology
by: Gemma Team, et al.
Published: (2024)
by: Gemma Team, et al.
Published: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints
by: Bologna, Federica, et al.
Published: (2025)
by: Bologna, Federica, et al.
Published: (2025)
AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
by: Matys, Piotr, et al.
Published: (2025)
by: Matys, Piotr, et al.
Published: (2025)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
by: Shelat, Shlok, et al.
Published: (2026)
by: Shelat, Shlok, et al.
Published: (2026)
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA
by: Baqar, Mohammad, et al.
Published: (2025)
by: Baqar, Mohammad, et al.
Published: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
REFIND at SemEval-2025 Task 3: Retrieval-Augmented Factuality Hallucination Detection in Large Language Models
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
Gemma 2: Improving Open Language Models at a Practical Size
by: Gemma Team, et al.
Published: (2024)
by: Gemma Team, et al.
Published: (2024)
EmbeddingGemma: Powerful and Lightweight Text Representations
by: Vera, Henrique Schechter, et al.
Published: (2025)
by: Vera, Henrique Schechter, et al.
Published: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
by: Lieberum, Tom, et al.
Published: (2024)
by: Lieberum, Tom, et al.
Published: (2024)
Investigating Hallucinations in Pruned Large Language Models for Abstractive Summarization
by: Chrysostomou, George, et al.
Published: (2023)
by: Chrysostomou, George, et al.
Published: (2023)
HausaNLP at SemEval-2025 Task 3: Towards a Fine-Grained Model-Aware Hallucination Detection
by: Bala, Maryam, et al.
Published: (2025)
by: Bala, Maryam, et al.
Published: (2025)
Similar Items
-
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025) -
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024) -
Halu-J: Critique-Based Hallucination Judge
by: Wang, Binjie, et al.
Published: (2024) -
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024) -
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
by: Chen, Kedi, et al.
Published: (2024)