Metric-Fair Prompting: Treating Similar Samples Similarly

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Jing, Shen, Jie, Niu, Xing, Zhang, Tong, Weiss, Jeremy
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915661209403392
author Wang, Jing
Shen, Jie
Niu, Xing
Zhang, Tong
Weiss, Jeremy
author_facet Wang, Jing
Shen, Jie
Niu, Xing
Zhang, Tong
Weiss, Jeremy
contents We introduce \emph{Metric-Fair Prompting}, a fairness-aware prompting framework that guides large language models (LLMs) to make decisions under metric-fairness constraints. In the application of multiple-choice medical question answering, each {(question, option)} pair is treated as a binary instance with label $+1$ (correct) or $-1$ (incorrect). To promote {individual fairness}~--~treating similar instances similarly~--~we compute question similarity using NLP embeddings and solve items in \emph{joint pairs of similar questions} rather than in isolation. The prompt enforces a global decision protocol: extract decisive clinical features, map each \((\text{question}, \text{option})\) to a score $f(x)$ that acts as confidence, and impose a Lipschitz-style constraint so that similar inputs receive similar scores and, hence, consistent outputs. Evaluated on the {MedQA (US)} benchmark, Metric-Fair Prompting is shown to improve performance over standard single-item prompting, demonstrating that fairness-guided, confidence-oriented reasoning can enhance LLM accuracy on high-stakes clinical multiple-choice questions.
format Preprint
id arxiv_https___arxiv_org_abs_2512_07608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Metric-Fair Prompting: Treating Similar Samples Similarly
Wang, Jing
Shen, Jie
Niu, Xing
Zhang, Tong
Weiss, Jeremy
Computation and Language
Artificial Intelligence
We introduce \emph{Metric-Fair Prompting}, a fairness-aware prompting framework that guides large language models (LLMs) to make decisions under metric-fairness constraints. In the application of multiple-choice medical question answering, each {(question, option)} pair is treated as a binary instance with label $+1$ (correct) or $-1$ (incorrect). To promote {individual fairness}~--~treating similar instances similarly~--~we compute question similarity using NLP embeddings and solve items in \emph{joint pairs of similar questions} rather than in isolation. The prompt enforces a global decision protocol: extract decisive clinical features, map each \((\text{question}, \text{option})\) to a score $f(x)$ that acts as confidence, and impose a Lipschitz-style constraint so that similar inputs receive similar scores and, hence, consistent outputs. Evaluated on the {MedQA (US)} benchmark, Metric-Fair Prompting is shown to improve performance over standard single-item prompting, demonstrating that fairness-guided, confidence-oriented reasoning can enhance LLM accuracy on high-stakes clinical multiple-choice questions.
title Metric-Fair Prompting: Treating Similar Samples Similarly
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.07608