Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Salla, Rohit Kumar, Saravanan, Manoj, Kota, Shrikar Reddy
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911346637930496
author Salla, Rohit Kumar
Saravanan, Manoj
Kota, Shrikar Reddy
author_facet Salla, Rohit Kumar
Saravanan, Manoj
Kota, Shrikar Reddy
contents Large Language Models (LLMs) like LLaMA, Mistral, and Gemma are increasingly used in decision-critical domains such as healthcare, law, and finance, yet their reliability remains uncertain. They often make overconfident errors, degrade under input shifts, and lack clear uncertainty estimates. Existing evaluations are fragmented, addressing only isolated aspects. We introduce the Composite Reliability Score (CRS), a unified framework that integrates calibration, robustness, and uncertainty quantification into a single interpretable metric. Through experiments on ten leading open-source LLMs across five QA datasets, we assess performance under baselines, perturbations, and calibration methods. CRS delivers stable model rankings, uncovers hidden failure modes missed by single metrics, and highlights that the most dependable systems balance accuracy, robustness, and calibrated uncertainty.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
Salla, Rohit Kumar
Saravanan, Manoj
Kota, Shrikar Reddy
Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
Large Language Models (LLMs) like LLaMA, Mistral, and Gemma are increasingly used in decision-critical domains such as healthcare, law, and finance, yet their reliability remains uncertain. They often make overconfident errors, degrade under input shifts, and lack clear uncertainty estimates. Existing evaluations are fragmented, addressing only isolated aspects. We introduce the Composite Reliability Score (CRS), a unified framework that integrates calibration, robustness, and uncertainty quantification into a single interpretable metric. Through experiments on ten leading open-source LLMs across five QA datasets, we assess performance under baselines, perturbations, and calibration methods. CRS delivers stable model rankings, uncovers hidden failure modes missed by single metrics, and highlights that the most dependable systems balance accuracy, robustness, and calibrated uncertainty.
title Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
I.2.7; I.2.6
url https://arxiv.org/abs/2512.24058