OLAPH: Improving Factuality in Biomedical Long-form Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeong, Minbyul, Hwang, Hyeon, Yoon, Chanwoong, Lee, Taewhoo, Kang, Jaewoo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913546442375168
author Jeong, Minbyul
Hwang, Hyeon
Yoon, Chanwoong
Lee, Taewhoo
Kang, Jaewoo
author_facet Jeong, Minbyul
Hwang, Hyeon
Yoon, Chanwoong
Lee, Taewhoo
Kang, Jaewoo
contents In the medical domain, numerous scenarios necessitate the long-form generation ability of large language models (LLMs). Specifically, when addressing patients' questions, it is essential that the model's response conveys factual claims, highlighting the need for an automated method to evaluate those claims. Thus, we introduce MedLFQA, a benchmark dataset reconstructed using long-form question-answering datasets related to the biomedical domain. We use MedLFQA to facilitate a cost-effective automatic evaluations of factuality. We also propose OLAPH, a simple and novel framework that utilizes cost-effective and multifaceted automatic evaluation to construct a synthetic preference set and answers questions in our preferred manner. Our framework leads us to train LLMs step-by-step to reduce hallucinations and include crucial medical claims. We highlight that, even on evaluation metrics not used during training, LLMs trained with our OLAPH framework demonstrate significant performance improvement in factuality. Our findings reveal that a 7B LLM trained with our OLAPH framework can provide long answers comparable to the medical experts' answers in terms of factuality. We believe that our work could shed light on gauging the long-text generation ability of LLMs in the medical domain. Our code and datasets are available.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12701
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OLAPH: Improving Factuality in Biomedical Long-form Question Answering
Jeong, Minbyul
Hwang, Hyeon
Yoon, Chanwoong
Lee, Taewhoo
Kang, Jaewoo
Computation and Language
Artificial Intelligence
In the medical domain, numerous scenarios necessitate the long-form generation ability of large language models (LLMs). Specifically, when addressing patients' questions, it is essential that the model's response conveys factual claims, highlighting the need for an automated method to evaluate those claims. Thus, we introduce MedLFQA, a benchmark dataset reconstructed using long-form question-answering datasets related to the biomedical domain. We use MedLFQA to facilitate a cost-effective automatic evaluations of factuality. We also propose OLAPH, a simple and novel framework that utilizes cost-effective and multifaceted automatic evaluation to construct a synthetic preference set and answers questions in our preferred manner. Our framework leads us to train LLMs step-by-step to reduce hallucinations and include crucial medical claims. We highlight that, even on evaluation metrics not used during training, LLMs trained with our OLAPH framework demonstrate significant performance improvement in factuality. Our findings reveal that a 7B LLM trained with our OLAPH framework can provide long answers comparable to the medical experts' answers in terms of factuality. We believe that our work could shed light on gauging the long-text generation ability of LLMs in the medical domain. Our code and datasets are available.
title OLAPH: Improving Factuality in Biomedical Long-form Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.12701