A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Flanagan, William, Das, Mukunda, Ramanayake, Rajitha, Maslekar, Swanuja, Mangipudi, Meghana, Choi, Joong Ho, Nair, Shruti, Bhusan, Shambhavi, Dulam, Sanjana, Pendharkar, Mouni, Singh, Nidhi, Doshi, Vashisth, Paresh, Sachi Shah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908596537655296
author Flanagan, William
Das, Mukunda
Ramanayake, Rajitha
Maslekar, Swanuja
Mangipudi, Meghana
Choi, Joong Ho
Nair, Shruti
Bhusan, Shambhavi
Dulam, Sanjana
Pendharkar, Mouni
Singh, Nidhi
Doshi, Vashisth
Paresh, Sachi Shah
author_facet Flanagan, William
Das, Mukunda
Ramanayake, Rajitha
Maslekar, Swanuja
Mangipudi, Meghana
Choi, Joong Ho
Nair, Shruti
Bhusan, Shambhavi
Dulam, Sanjana
Pendharkar, Mouni
Singh, Nidhi
Doshi, Vashisth
Paresh, Sachi Shah
contents As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to GenAI workloads and are often supplemented using Subject Matter Expert (SME) Evaluation. Even in this combination, many projects fail to account for various unique risks present in choosing specific metrics. Additionally, many widespread benchmarks created by foundational research labs and educational institutions fail to generalize to industrial use. This paper explains these challenges and provides a Risk Assessment Framework to allow for better application of SME and machine learning Metrics
format Preprint
id arxiv_https___arxiv_org_abs_2510_13524
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain
Flanagan, William
Das, Mukunda
Ramanayake, Rajitha
Maslekar, Swanuja
Mangipudi, Meghana
Choi, Joong Ho
Nair, Shruti
Bhusan, Shambhavi
Dulam, Sanjana
Pendharkar, Mouni
Singh, Nidhi
Doshi, Vashisth
Paresh, Sachi Shah
Artificial Intelligence
As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to GenAI workloads and are often supplemented using Subject Matter Expert (SME) Evaluation. Even in this combination, many projects fail to account for various unique risks present in choosing specific metrics. Additionally, many widespread benchmarks created by foundational research labs and educational institutions fail to generalize to industrial use. This paper explains these challenges and provides a Risk Assessment Framework to allow for better application of SME and machine learning Metrics
title A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain
topic Artificial Intelligence
url https://arxiv.org/abs/2510.13524