HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Anaokar, Spandan, Ganatra, Shrey, Kashid, Harshvivek, Bhattacharyya, Swapnil, Nair, Shruti, Sekhar, Reshma, Manohar, Siddharth, Hemrajani, Rahul, Bhattacharyya, Pushpak
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909839248064512
author Anaokar, Spandan
Ganatra, Shrey
Kashid, Harshvivek
Bhattacharyya, Swapnil
Nair, Shruti
Sekhar, Reshma
Manohar, Siddharth
Hemrajani, Rahul
Bhattacharyya, Pushpak
author_facet Anaokar, Spandan
Ganatra, Shrey
Kashid, Harshvivek
Bhattacharyya, Swapnil
Nair, Shruti
Sekhar, Reshma
Manohar, Siddharth
Hemrajani, Rahul
Bhattacharyya, Pushpak
contents Large Language Models (LLMs) are widely used in industry but remain prone to hallucinations, limiting their reliability in critical applications. This work addresses hallucination reduction in consumer grievance chatbots built using LLaMA 3.1 8B Instruct, a compact model frequently used in industry. We develop HalluDetect, an LLM-based hallucination detection system that achieves an F1 score of 68.92% outperforming baseline detectors by 22.47%. Benchmarking five hallucination mitigation architectures, we find that out of them, AgentBot minimizes hallucinations to 0.4159 per turn while maintaining the highest token accuracy (96.13%), making it the most effective mitigation strategy. Our findings provide a scalable framework for hallucination mitigation, demonstrating that optimized inference strategies can significantly improve factual accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
Anaokar, Spandan
Ganatra, Shrey
Kashid, Harshvivek
Bhattacharyya, Swapnil
Nair, Shruti
Sekhar, Reshma
Manohar, Siddharth
Hemrajani, Rahul
Bhattacharyya, Pushpak
Computation and Language
Large Language Models (LLMs) are widely used in industry but remain prone to hallucinations, limiting their reliability in critical applications. This work addresses hallucination reduction in consumer grievance chatbots built using LLaMA 3.1 8B Instruct, a compact model frequently used in industry. We develop HalluDetect, an LLM-based hallucination detection system that achieves an F1 score of 68.92% outperforming baseline detectors by 22.47%. Benchmarking five hallucination mitigation architectures, we find that out of them, AgentBot minimizes hallucinations to 0.4159 per turn while maintaining the highest token accuracy (96.13%), making it the most effective mitigation strategy. Our findings provide a scalable framework for hallucination mitigation, demonstrating that optimized inference strategies can significantly improve factual accuracy.
title HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain
topic Computation and Language
url https://arxiv.org/abs/2509.11619