WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maurya, Sneha, Saboo, Pragya, Kumar, Girish
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912993212628992
author Maurya, Sneha
Saboo, Pragya
Kumar, Girish
author_facet Maurya, Sneha
Saboo, Pragya
Kumar, Girish
contents Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios across 10 women's health topics, designed to expose clinically meaningful failure modes including outdated guidelines, unsafe omissions, dosing errors, and equity-related blind spots. We evaluate 22 models using a 23-criterion rubric spanning clinical accuracy, completeness, safety, communication quality, instruction following, equity, uncertainty handling, and guideline adherence, with safety-weighted penalties and server-side score recalculation. Across 3,102 attempted responses (3,100 scored), no model mean performance exceeds 75 percent; the best model reaches 72.1 percent. Even top models show low fully correct rates and substantial variation in harm rates. Inter-rater reliability is moderate at the response label level but high for model ranking, supporting WHBench utility for comparative system evaluation while highlighting the need for expert oversight in clinical deployment. WHBench provides a public, failure-mode-aware benchmark to track safer and more equitable progress in womens health AI.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00024
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
Maurya, Sneha
Saboo, Pragya
Kumar, Girish
Computation and Language
Artificial Intelligence
Computers and Society
Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Women's Health Benchmark (WHBench), a targeted evaluation suite of 47 expert-crafted scenarios across 10 women's health topics, designed to expose clinically meaningful failure modes including outdated guidelines, unsafe omissions, dosing errors, and equity-related blind spots. We evaluate 22 models using a 23-criterion rubric spanning clinical accuracy, completeness, safety, communication quality, instruction following, equity, uncertainty handling, and guideline adherence, with safety-weighted penalties and server-side score recalculation. Across 3,102 attempted responses (3,100 scored), no model mean performance exceeds 75 percent; the best model reaches 72.1 percent. Even top models show low fully correct rates and substantial variation in harm rates. Inter-rater reliability is moderate at the response label level but high for model ranking, supporting WHBench utility for comparative system evaluation while highlighting the need for expert oversight in clinical deployment. WHBench provides a public, failure-mode-aware benchmark to track safer and more equitable progress in womens health AI.
title WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2604.00024