A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908596537655296 |
|---|---|
| author | Flanagan, William Das, Mukunda Ramanayake, Rajitha Maslekar, Swanuja Mangipudi, Meghana Choi, Joong Ho Nair, Shruti Bhusan, Shambhavi Dulam, Sanjana Pendharkar, Mouni Singh, Nidhi Doshi, Vashisth Paresh, Sachi Shah |
| author_facet | Flanagan, William Das, Mukunda Ramanayake, Rajitha Maslekar, Swanuja Mangipudi, Meghana Choi, Joong Ho Nair, Shruti Bhusan, Shambhavi Dulam, Sanjana Pendharkar, Mouni Singh, Nidhi Doshi, Vashisth Paresh, Sachi Shah |
| contents | As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to GenAI workloads and are often supplemented using Subject Matter Expert (SME) Evaluation. Even in this combination, many projects fail to account for various unique risks present in choosing specific metrics. Additionally, many widespread benchmarks created by foundational research labs and educational institutions fail to generalize to industrial use. This paper explains these challenges and provides a Risk Assessment Framework to allow for better application of SME and machine learning Metrics |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13524 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain Flanagan, William Das, Mukunda Ramanayake, Rajitha Maslekar, Swanuja Mangipudi, Meghana Choi, Joong Ho Nair, Shruti Bhusan, Shambhavi Dulam, Sanjana Pendharkar, Mouni Singh, Nidhi Doshi, Vashisth Paresh, Sachi Shah Artificial Intelligence As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical machine learning metrics can oftentimes fail to generalize to GenAI workloads and are often supplemented using Subject Matter Expert (SME) Evaluation. Even in this combination, many projects fail to account for various unique risks present in choosing specific metrics. Additionally, many widespread benchmarks created by foundational research labs and educational institutions fail to generalize to industrial use. This paper explains these challenges and provides a Risk Assessment Framework to allow for better application of SME and machine learning Metrics |
| title | A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2510.13524 |