PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Salimian, Sina, Uddin, Gias, Raza, Shaina, Leung, Henry
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909893643993088
author Salimian, Sina
Uddin, Gias
Raza, Shaina
Leung, Henry
author_facet Salimian, Sina
Uddin, Gias
Raza, Shaina
Leung, Henry
contents Zero-shot LLMs are now also used for textual classification tasks, e.g., sentiment and bias detection in a sentence or article. However, their performance can be suboptimal in such data annotation tasks. We introduce a novel technique that evaluates an LLM's confidence for classifying a textual input by leveraging Metamorphic Relations (MRs). The MRs generate semantically equivalent yet textually divergent versions of the input. Following the principles of Metamorphic Testing (MT), the mutated versions are expected to have annotation labels similar to the input. By analyzing the consistency of an LLM's responses across these variations, we compute a perceived confidence score (PCS) based on the frequency of the predicted labels. PCS can be used for both single and multiple LLM settings (e.g., when multiple LLMs are vetted in a majority-voting setup). Empirical evaluation shows that our PCS-based approach improves the performance of zero-shot LLMs by 9.3% in textual classification tasks. When multiple LLMs are used in a majority-voting setup, we obtain a performance boost of 5.8% with PCS.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
Salimian, Sina
Uddin, Gias
Raza, Shaina
Leung, Henry
Computation and Language
Machine Learning
Zero-shot LLMs are now also used for textual classification tasks, e.g., sentiment and bias detection in a sentence or article. However, their performance can be suboptimal in such data annotation tasks. We introduce a novel technique that evaluates an LLM's confidence for classifying a textual input by leveraging Metamorphic Relations (MRs). The MRs generate semantically equivalent yet textually divergent versions of the input. Following the principles of Metamorphic Testing (MT), the mutated versions are expected to have annotation labels similar to the input. By analyzing the consistency of an LLM's responses across these variations, we compute a perceived confidence score (PCS) based on the frequency of the predicted labels. PCS can be used for both single and multiple LLM settings (e.g., when multiple LLMs are vetted in a majority-voting setup). Empirical evaluation shows that our PCS-based approach improves the performance of zero-shot LLMs by 9.3% in textual classification tasks. When multiple LLMs are used in a majority-voting setup, we obtain a performance boost of 5.8% with PCS.
title PCS: Perceived Confidence Scoring of Black Box LLMs with Metamorphic Relations
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.07186