HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Boshui, Fan, Zhaoxin, Wang, Ke, Leng, Zhiying, Wu, Faguo, Zheng, Hongwei, Sun, Yifan, Wu, Wenjun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lyapunov Probes for Hallucination Detection in Large Foundation Models
by: Luan, Bozhi, et al.
Published: (2026)
by: Luan, Bozhi, et al.
Published: (2026)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)
by: Alansari, Aisha, et al.
Published: (2026)
HalluZig: Hallucination Detection using Zigzag Persistence
by: Samaga, Shreyas N., et al.
Published: (2026)
by: Samaga, Shreyas N., et al.
Published: (2026)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
Sparse Auto-Encoders and Holism about Large Language Models
by: Grindrod, Jumbly
Published: (2026)
by: Grindrod, Jumbly
Published: (2026)
HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
by: Urlana, Ashok, et al.
Published: (2025)
by: Urlana, Ashok, et al.
Published: (2025)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
by: Yeh, Min-Hsuan, et al.
Published: (2025)
by: Yeh, Min-Hsuan, et al.
Published: (2025)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
by: Anaokar, Spandan, et al.
Published: (2025)
by: Anaokar, Spandan, et al.
Published: (2025)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
by: Jing, Yi, et al.
Published: (2025)
by: Jing, Yi, et al.
Published: (2025)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025)
by: Dasgupta, Sharanya, et al.
Published: (2025)
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild
by: Zhu, Zhiying, et al.
Published: (2024)
by: Zhu, Zhiying, et al.
Published: (2024)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
by: Luo, Wen, et al.
Published: (2024)
by: Luo, Wen, et al.
Published: (2024)
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
by: Adib, Shefayat E Shams, et al.
Published: (2026)
by: Adib, Shefayat E Shams, et al.
Published: (2026)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
by: Noël, Valentin, et al.
Published: (2025)
by: Noël, Valentin, et al.
Published: (2025)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
by: Hosseini, Faezeh, et al.
Published: (2026)
by: Hosseini, Faezeh, et al.
Published: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
by: Li, Tianyi, et al.
Published: (2024)
by: Li, Tianyi, et al.
Published: (2024)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
by: Zhao, Yaxin, et al.
Published: (2025)
by: Zhao, Yaxin, et al.
Published: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
by: Ghilardi, Davide, et al.
Published: (2024)
by: Ghilardi, Davide, et al.
Published: (2024)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026)
by: Liu, Emmy, et al.
Published: (2026)
Exploring Task Performance with Interpretable Models via Sparse Auto-Encoders
by: Wang, Shun, et al.
Published: (2025)
by: Wang, Shun, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
by: Ridder, Fabian, et al.
Published: (2024)
by: Ridder, Fabian, et al.
Published: (2024)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
by: Nath, Sujoy, et al.
Published: (2025)
by: Nath, Sujoy, et al.
Published: (2025)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
by: Wei, Jiacheng, et al.
Published: (2025)
by: Wei, Jiacheng, et al.
Published: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
by: Abdallah, Mohamed A., et al.
Published: (2025)
by: Abdallah, Mohamed A., et al.
Published: (2025)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
by: Bergeron, Loris, et al.
Published: (2025)
by: Bergeron, Loris, et al.
Published: (2025)
AutoHall: Automated Factuality Hallucination Dataset Generation for Large Language Models
by: Cao, Zouying, et al.
Published: (2023)
by: Cao, Zouying, et al.
Published: (2023)
Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models
by: Zhang, Hanzhi, et al.
Published: (2025)
by: Zhang, Hanzhi, et al.
Published: (2025)
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
by: Aparin, Georgii, et al.
Published: (2026)
by: Aparin, Georgii, et al.
Published: (2026)
AlignSAE: Concept-Aligned Sparse Autoencoders
by: Yang, Minglai, et al.
Published: (2025)
by: Yang, Minglai, et al.
Published: (2025)
Similar Items
-
Lyapunov Probes for Hallucination Detection in Large Foundation Models
by: Luan, Bozhi, et al.
Published: (2026) -
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026) -
HalluZig: Hallucination Detection using Zigzag Persistence
by: Samaga, Shreyas N., et al.
Published: (2026) -
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025) -
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)