Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Xinping, Yu, Jindi, Liu, Zhenyu, Wang, Jifang, Li, Dongfang, Chen, Yibin, Hu, Baotian, Zhang, Min
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917802008379392
author Zhao, Xinping
Yu, Jindi
Liu, Zhenyu
Wang, Jifang
Li, Dongfang
Chen, Yibin
Hu, Baotian
Zhang, Min
author_facet Zhao, Xinping
Yu, Jindi
Liu, Zhenyu
Wang, Jifang
Li, Dongfang
Chen, Yibin
Hu, Baotian
Zhang, Min
contents As we all know, hallucinations prevail in Large Language Models (LLMs), where the generated content is coherent but factually incorrect, which inflicts a heavy blow on the widespread application of LLMs. Previous studies have shown that LLMs could confidently state non-existent facts rather than answering ``I don't know''. Therefore, it is necessary to resort to external knowledge to detect and correct the hallucinated content. Since manual detection and correction of factual errors is labor-intensive, developing an automatic end-to-end hallucination-checking approach is indeed a needful thing. To this end, we present Medico, a Multi-source evidence fusion enhanced hallucination detection and correction framework. It fuses diverse evidence from multiple sources, detects whether the generated content contains factual errors, provides the rationale behind the judgment, and iteratively revises the hallucinated content. Experimental results on evidence retrieval (0.964 HR@5, 0.908 MRR@5), hallucination detection (0.927-0.951 F1), and hallucination correction (0.973-0.979 approval rate) manifest the great potential of Medico. A video demo of Medico can be found at https://youtu.be/RtsO6CSesBI.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10408
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
Zhao, Xinping
Yu, Jindi
Liu, Zhenyu
Wang, Jifang
Li, Dongfang
Chen, Yibin
Hu, Baotian
Zhang, Min
Computation and Language
Information Retrieval
As we all know, hallucinations prevail in Large Language Models (LLMs), where the generated content is coherent but factually incorrect, which inflicts a heavy blow on the widespread application of LLMs. Previous studies have shown that LLMs could confidently state non-existent facts rather than answering ``I don't know''. Therefore, it is necessary to resort to external knowledge to detect and correct the hallucinated content. Since manual detection and correction of factual errors is labor-intensive, developing an automatic end-to-end hallucination-checking approach is indeed a needful thing. To this end, we present Medico, a Multi-source evidence fusion enhanced hallucination detection and correction framework. It fuses diverse evidence from multiple sources, detects whether the generated content contains factual errors, provides the rationale behind the judgment, and iteratively revises the hallucinated content. Experimental results on evidence retrieval (0.964 HR@5, 0.908 MRR@5), hallucination detection (0.927-0.951 F1), and hallucination correction (0.973-0.979 approval rate) manifest the great potential of Medico. A video demo of Medico can be found at https://youtu.be/RtsO6CSesBI.
title Medico: Towards Hallucination Detection and Correction with Multi-source Evidence Fusion
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2410.10408