Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ozaki, Shintaro, Kato, Yuta, Feng, Siyuan, Tomita, Masayo, Hayashi, Kazuki, Hashimoto, Wataru, Obara, Ryoma, Oyamada, Masafumi, Hayashi, Katsuhiko, Kamigaito, Hidetaka, Watanabe, Taro
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911110652755968
author Ozaki, Shintaro
Kato, Yuta
Feng, Siyuan
Tomita, Masayo
Hayashi, Kazuki
Hashimoto, Wataru
Obara, Ryoma
Oyamada, Masafumi
Hayashi, Katsuhiko
Kamigaito, Hidetaka
Watanabe, Taro
author_facet Ozaki, Shintaro
Kato, Yuta
Feng, Siyuan
Tomita, Masayo
Hayashi, Kazuki
Hashimoto, Wataru
Obara, Ryoma
Oyamada, Masafumi
Hayashi, Katsuhiko
Kamigaito, Hidetaka
Watanabe, Taro
contents Retrieval Augmented Generation (RAG) complements the knowledge of Large Language Models (LLMs) by leveraging external information to enhance response accuracy for queries. This approach is widely applied in several fields by taking its advantage of injecting the most up-to-date information, and researchers are focusing on understanding and improving this aspect to unlock the full potential of RAG in such high-stakes applications. However, despite the potential of RAG to address these needs, the mechanisms behind the confidence levels of its outputs remain underexplored. Our study focuses on the impact of RAG, specifically examining whether RAG improves the confidence of LLM outputs in the medical domain. We conduct this analysis across various configurations and models. We evaluate confidence by treating the model's predicted probability as its output and calculating several evaluation metrics which include calibration error method, entropy, the best probability, and accuracy. Experimental results across multiple datasets confirmed that certain models possess the capability to judge for themselves whether an inserted document relates to the correct answer. These results suggest that evaluating models based on their output probabilities determine whether they function as generators in the RAG framework. Our approach allows us to evaluate whether the models handle retrieved documents.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain
Ozaki, Shintaro
Kato, Yuta
Feng, Siyuan
Tomita, Masayo
Hayashi, Kazuki
Hashimoto, Wataru
Obara, Ryoma
Oyamada, Masafumi
Hayashi, Katsuhiko
Kamigaito, Hidetaka
Watanabe, Taro
Computation and Language
Retrieval Augmented Generation (RAG) complements the knowledge of Large Language Models (LLMs) by leveraging external information to enhance response accuracy for queries. This approach is widely applied in several fields by taking its advantage of injecting the most up-to-date information, and researchers are focusing on understanding and improving this aspect to unlock the full potential of RAG in such high-stakes applications. However, despite the potential of RAG to address these needs, the mechanisms behind the confidence levels of its outputs remain underexplored. Our study focuses on the impact of RAG, specifically examining whether RAG improves the confidence of LLM outputs in the medical domain. We conduct this analysis across various configurations and models. We evaluate confidence by treating the model's predicted probability as its output and calculating several evaluation metrics which include calibration error method, entropy, the best probability, and accuracy. Experimental results across multiple datasets confirmed that certain models possess the capability to judge for themselves whether an inserted document relates to the correct answer. These results suggest that evaluating models based on their output probabilities determine whether they function as generators in the RAG framework. Our approach allows us to evaluate whether the models handle retrieved documents.
title Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain
topic Computation and Language
url https://arxiv.org/abs/2412.20309