Saved in:
Bibliographic Details
Main Authors: Gondara, Lovedeep, Simkin, Jonathan
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.16702
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • We present an audit mechanism for language models, with a focus on models deployed in the healthcare setting. Our proposed mechanism takes inspiration from clinical trial design where we posit the language model audit as a single blind equivalence trial, with the comparison of interest being the subject matter experts. We show that using our proposed method, we can follow principled sample size and power calculations, leading to the requirement of sampling minimum number of records while maintaining the audit integrity and statistical soundness. Finally, we provide a real-world example of the audit used in a production environment in a large-scale public health network.