How to Choose a Threshold for an Evaluation Metric for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sarmah, Bhaskarjit, Li, Mingshu, Lyu, Jingrao, Frank, Sebastian, Castellanos, Nathalia, Pasquali, Stefano, Mehta, Dhagash
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912159119704064
author Sarmah, Bhaskarjit
Li, Mingshu
Lyu, Jingrao
Frank, Sebastian
Castellanos, Nathalia
Pasquali, Stefano
Mehta, Dhagash
author_facet Sarmah, Bhaskarjit
Li, Mingshu
Lyu, Jingrao
Frank, Sebastian
Castellanos, Nathalia
Pasquali, Stefano
Mehta, Dhagash
contents To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a methodology to identify a robust threshold on these metrics even though there are many serious implications of an incorrect choice of the thresholds during deployment of the LLMs. Translating the traditional model risk management (MRM) guidelines within regulated industries such as the financial industry, we propose a step-by-step recipe for picking a threshold for a given LLM evaluation metric. We emphasize that such a methodology should start with identifying the risks of the LLM application under consideration and risk tolerance of the stakeholders. We then propose concrete and statistically rigorous procedures to determine a threshold for the given LLM evaluation metric using available ground-truth data. As a concrete example to demonstrate the proposed methodology at work, we employ it on the Faithfulness metric, as implemented in various publicly available libraries, using the publicly available HaluBench dataset. We also lay a foundation for creating systematic approaches to select thresholds, not only for LLMs but for any GenAI applications.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12148
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle How to Choose a Threshold for an Evaluation Metric for Large Language Models
Sarmah, Bhaskarjit
Li, Mingshu
Lyu, Jingrao
Frank, Sebastian
Castellanos, Nathalia
Pasquali, Stefano
Mehta, Dhagash
Machine Learning
Computation and Language
Statistical Finance
Applications
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a methodology to identify a robust threshold on these metrics even though there are many serious implications of an incorrect choice of the thresholds during deployment of the LLMs. Translating the traditional model risk management (MRM) guidelines within regulated industries such as the financial industry, we propose a step-by-step recipe for picking a threshold for a given LLM evaluation metric. We emphasize that such a methodology should start with identifying the risks of the LLM application under consideration and risk tolerance of the stakeholders. We then propose concrete and statistically rigorous procedures to determine a threshold for the given LLM evaluation metric using available ground-truth data. As a concrete example to demonstrate the proposed methodology at work, we employ it on the Faithfulness metric, as implemented in various publicly available libraries, using the publicly available HaluBench dataset. We also lay a foundation for creating systematic approaches to select thresholds, not only for LLMs but for any GenAI applications.
title How to Choose a Threshold for an Evaluation Metric for Large Language Models
topic Machine Learning
Computation and Language
Statistical Finance
Applications
url https://arxiv.org/abs/2412.12148