Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nahian, Md Sultan Al, Kavuluru, Ramakanth
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917727739838464
author Nahian, Md Sultan Al
Kavuluru, Ramakanth
author_facet Nahian, Md Sultan Al
Kavuluru, Ramakanth
contents Extractive question answering over clinical text is a crucial need to help deal with the deluge of clinical text generated in hospitals. While encoder models (e.g., BERT) have been popular for this reading comprehension task, recently encoder-decoder models (e.g., T5) are on the rise. There is also the emergence of preference optimization techniques to align decoder-only LLMs with human preferences. In this paper, we combine encoder-decoder models with the direct preference optimization (DPO) method to improve over prior state of the art for the RadQA radiology question answering task by 12-15 F1 points. To the best of our knowledge, this effort is the first to show that DPO method also works for reading comprehension via novel heuristics to generate preference data without human inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14000
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization
Nahian, Md Sultan Al
Kavuluru, Ramakanth
Information Retrieval
Computation and Language
Machine Learning
Extractive question answering over clinical text is a crucial need to help deal with the deluge of clinical text generated in hospitals. While encoder models (e.g., BERT) have been popular for this reading comprehension task, recently encoder-decoder models (e.g., T5) are on the rise. There is also the emergence of preference optimization techniques to align decoder-only LLMs with human preferences. In this paper, we combine encoder-decoder models with the direct preference optimization (DPO) method to improve over prior state of the art for the RadQA radiology question answering task by 12-15 F1 points. To the best of our knowledge, this effort is the first to show that DPO method also works for reading comprehension via novel heuristics to generate preference data without human inputs.
title Clinical Reading Comprehension with Encoder-Decoder Models Enhanced by Direct Preference Optimization
topic Information Retrieval
Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.14000