Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De la Iglesia, Iker, Goenaga, Iakes, Ramirez-Romero, Johanna, Villa-Gonzalez, Jose Maria, Goikoetxea, Josu, Barrena, Ander
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909330965528576
author De la Iglesia, Iker
Goenaga, Iakes
Ramirez-Romero, Johanna
Villa-Gonzalez, Jose Maria
Goikoetxea, Josu
Barrena, Ander
author_facet De la Iglesia, Iker
Goenaga, Iakes
Ramirez-Romero, Johanna
Villa-Gonzalez, Jose Maria
Goikoetxea, Josu
Barrena, Ander
contents Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM-generated medical explanatory arguments, relying on Proxy Tasks and rankings to closely align results with human evaluation criteria, overcoming the biases typically seen in LLMs used as judges. We demonstrate that the proposed evaluators are robust against adversarial attacks, including the assessment of non-argumentative text. Additionally, the human-crafted arguments needed to train the evaluators are minimized to just one example per Proxy Task. By examining multiple LLM-generated arguments, we establish a methodology for determining whether a Proxy Task is suitable for evaluating LLM-generated medical explanatory arguments, requiring only five examples and two human experts.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20565
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments
De la Iglesia, Iker
Goenaga, Iakes
Ramirez-Romero, Johanna
Villa-Gonzalez, Jose Maria
Goikoetxea, Josu
Barrena, Ander
Computation and Language
Evaluating LLM-generated text has become a key challenge, especially in domain-specific contexts like the medical field. This work introduces a novel evaluation methodology for LLM-generated medical explanatory arguments, relying on Proxy Tasks and rankings to closely align results with human evaluation criteria, overcoming the biases typically seen in LLMs used as judges. We demonstrate that the proposed evaluators are robust against adversarial attacks, including the assessment of non-argumentative text. Additionally, the human-crafted arguments needed to train the evaluators are minimized to just one example per Proxy Task. By examining multiple LLM-generated arguments, we establish a methodology for determining whether a Proxy Task is suitable for evaluating LLM-generated medical explanatory arguments, requiring only five examples and two human experts.
title Ranking Over Scoring: Towards Reliable and Robust Automated Evaluation of LLM-Generated Medical Explanatory Arguments
topic Computation and Language
url https://arxiv.org/abs/2409.20565