HalluCounter: Reference-free LLM Hallucination Detection in the Wild!
Fuente:
arXiv
Saved in:
| Main Authors: | Urlana, Ashok, Kanumolu, Gopichand, Kumar, Charaka Vinayak, Garlapati, Bala Mallikarjunarao, Mishra, Rahul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
by: Kumar, Charaka Vinayak, et al.
Published: (2025)
Agent Ideate: A Framework for Product Idea Generation from Patents Using Agentic AI
by: Kanumolu, Gopichand, et al.
Published: (2025)
by: Kanumolu, Gopichand, et al.
Published: (2025)
LLMs with Industrial Lens: Deciphering the Challenges and Prospects -- A Survey
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
No Size Fits All: The Perils and Pitfalls of Leveraging LLMs Vary with Company Size
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
by: P, Dinesh Srivasthav, et al.
Published: (2026)
by: P, Dinesh Srivasthav, et al.
Published: (2026)
Cyber for AI at SemEval-2025 Task 4: Forgotten but Not Lost: The Balancing Act of Selective Unlearning in Large Language Models
by: P, Dinesh Srivasthav, et al.
Published: (2025)
by: P, Dinesh Srivasthav, et al.
Published: (2025)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
by: Bala, Abhinaba, et al.
Published: (2024)
by: Bala, Abhinaba, et al.
Published: (2024)
TeClass: A Human-Annotated Relevance-based Headline Classification and Generation Dataset for Telugu
by: Kanumolu, Gopichand, et al.
Published: (2024)
by: Kanumolu, Gopichand, et al.
Published: (2024)
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects -- A Survey
by: Urlana, Ashok, et al.
Published: (2023)
by: Urlana, Ashok, et al.
Published: (2023)
LimGen: Probing the LLMs for Generating Suggestive Limitations of Research Papers
by: Faizullah, Abdur Rahman Bin Md, et al.
Published: (2024)
by: Faizullah, Abdur Rahman Bin Md, et al.
Published: (2024)
HalluLens: LLM Hallucination Benchmark
by: Bang, Yejin, et al.
Published: (2025)
by: Bang, Yejin, et al.
Published: (2025)
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
by: Anaokar, Spandan, et al.
Published: (2025)
by: Anaokar, Spandan, et al.
Published: (2025)
HalluZig: Hallucination Detection using Zigzag Persistence
by: Samaga, Shreyas N., et al.
Published: (2026)
by: Samaga, Shreyas N., et al.
Published: (2026)
HalluCana: Fixing LLM Hallucination with A Canary Lookahead
by: Li, Tianyi, et al.
Published: (2024)
by: Li, Tianyi, et al.
Published: (2024)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026)
by: Liu, Emmy, et al.
Published: (2026)
HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
by: Yeh, Min-Hsuan, et al.
Published: (2025)
by: Yeh, Min-Hsuan, et al.
Published: (2025)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
by: Ridder, Fabian, et al.
Published: (2024)
by: Ridder, Fabian, et al.
Published: (2024)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
by: Abdaljalil, Samir, et al.
Published: (2025)
by: Abdaljalil, Samir, et al.
Published: (2025)
HalluClean: A Unified Framework to Combat Hallucinations in LLMs
by: Zhao, Yaxin, et al.
Published: (2025)
by: Zhao, Yaxin, et al.
Published: (2025)
HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
by: Dasgupta, Sharanya, et al.
Published: (2025)
by: Dasgupta, Sharanya, et al.
Published: (2025)
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
by: Kumar, Mahesh, et al.
Published: (2026)
by: Kumar, Mahesh, et al.
Published: (2026)
HalluScore: Large Language Model Hallucination Question Answering Benchmark
by: Alansari, Aisha, et al.
Published: (2026)
by: Alansari, Aisha, et al.
Published: (2026)
HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders
by: Chen, Boshui, et al.
Published: (2026)
by: Chen, Boshui, et al.
Published: (2026)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
by: Fan, Dongyang, et al.
Published: (2026)
by: Fan, Dongyang, et al.
Published: (2026)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection
by: Emery, Deanna, et al.
Published: (2025)
by: Emery, Deanna, et al.
Published: (2025)
MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models
by: Pandit, Shrey, et al.
Published: (2025)
by: Pandit, Shrey, et al.
Published: (2025)
Reference-free Hallucination Detection for Large Vision-Language Models
by: Li, Qing, et al.
Published: (2024)
by: Li, Qing, et al.
Published: (2024)
AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs
by: Alansari, Aisha, et al.
Published: (2025)
by: Alansari, Aisha, et al.
Published: (2025)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
by: Hosseini, Faezeh, et al.
Published: (2026)
by: Hosseini, Faezeh, et al.
Published: (2026)
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
by: Luo, Wen, et al.
Published: (2024)
by: Luo, Wen, et al.
Published: (2024)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
by: Noël, Valentin, et al.
Published: (2025)
by: Noël, Valentin, et al.
Published: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
by: Cherif, Ahmed
Published: (2026)
by: Cherif, Ahmed
Published: (2026)
HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
by: Bergeron, Loris, et al.
Published: (2025)
by: Bergeron, Loris, et al.
Published: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
by: Abdallah, Mohamed A., et al.
Published: (2025)
by: Abdallah, Mohamed A., et al.
Published: (2025)
HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
by: Sakai, Yusuke, et al.
Published: (2026)
by: Sakai, Yusuke, et al.
Published: (2026)
Similar Items
-
No LLM is Free From Bias: A Comprehensive Study of Bias Evaluation in Large Language Models
by: Kumar, Charaka Vinayak, et al.
Published: (2025) -
Agent Ideate: A Framework for Product Idea Generation from Patents Using Agentic AI
by: Kanumolu, Gopichand, et al.
Published: (2025) -
LLMs with Industrial Lens: Deciphering the Challenges and Prospects -- A Survey
by: Urlana, Ashok, et al.
Published: (2024) -
No Size Fits All: The Perils and Pitfalls of Leveraging LLMs Vary with Company Size
by: Urlana, Ashok, et al.
Published: (2024) -
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
by: Urlana, Ashok, et al.
Published: (2024)