Measuring short-form factuality in large language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Jason, Karina, Nguyen, Chung, Hyung Won, Jiao, Yunxin Joy, Papay, Spencer, Glaese, Amelia, Schulman, John, Fedus, William
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912108385402880
author Wei, Jason
Karina, Nguyen
Chung, Hyung Won
Jiao, Yunxin Joy
Papay, Spencer
Glaese, Amelia
Schulman, John
Fedus, William
author_facet Wei, Jason
Karina, Nguyen
Chung, Hyung Won
Jiao, Yunxin Joy
Papay, Spencer
Glaese, Amelia
Schulman, John
Fedus, William
contents We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First, SimpleQA is challenging, as it is adversarially collected against GPT-4 responses. Second, responses are easy to grade, because questions are created such that there exists only a single, indisputable answer. Each answer in SimpleQA is graded as either correct, incorrect, or not attempted. A model with ideal behavior would get as many questions correct as possible while not attempting the questions for which it is not confident it knows the correct answer. SimpleQA is a simple, targeted evaluation for whether models "know what they know," and our hope is that this benchmark will remain relevant for the next few generations of frontier models. SimpleQA can be found at https://github.com/openai/simple-evals.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04368
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Measuring short-form factuality in large language models
Wei, Jason
Karina, Nguyen
Chung, Hyung Won
Jiao, Yunxin Joy
Papay, Spencer
Glaese, Amelia
Schulman, John
Fedus, William
Computation and Language
We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First, SimpleQA is challenging, as it is adversarially collected against GPT-4 responses. Second, responses are easy to grade, because questions are created such that there exists only a single, indisputable answer. Each answer in SimpleQA is graded as either correct, incorrect, or not attempted. A model with ideal behavior would get as many questions correct as possible while not attempting the questions for which it is not confident it knows the correct answer. SimpleQA is a simple, targeted evaluation for whether models "know what they know," and our hope is that this benchmark will remain relevant for the next few generations of frontier models. SimpleQA can be found at https://github.com/openai/simple-evals.
title Measuring short-form factuality in large language models
topic Computation and Language
url https://arxiv.org/abs/2411.04368