Generating Benchmarks for Factuality Evaluation of Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Muhlgay, Dor, Ram, Ori, Magar, Inbal, Levine, Yoav, Ratner, Nir, Belinkov, Yonatan, Abend, Omri, Leyton-Brown, Kevin, Shashua, Amnon, Shoham, Yoav
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916113677287424
author Muhlgay, Dor
Ram, Ori
Magar, Inbal
Levine, Yoav
Ratner, Nir
Belinkov, Yonatan
Abend, Omri
Leyton-Brown, Kevin
Shashua, Amnon
Shoham, Yoav
author_facet Muhlgay, Dor
Ram, Ori
Magar, Inbal
Levine, Yoav
Ratner, Nir
Belinkov, Yonatan
Abend, Omri
Leyton-Brown, Kevin
Shashua, Amnon
Shoham, Yoav
contents Before deploying a language model (LM) within a given domain, it is important to measure its tendency to generate factually incorrect information in that domain. Existing methods for factuality evaluation of LLM generation focus on facts sampled from the LM itself, and thus do not control the set of evaluated facts and might under-represent domain specific or rare facts. We propose FACTOR: Factual Assessment via Corpus TransfORmation, a scalable approach for evaluating LM factuality. FACTOR automatically transforms a factual corpus of interest into a benchmark evaluating an LM's propensity to generate true facts from the corpus vs. similar but incorrect statements. We use our framework to create three benchmarks: Wiki-FACTOR, News-FACTOR and Expert-FACTOR. We show that: (i) our benchmark scores increase with model size and improve when the LM is augmented with retrieval; (ii) benchmark score and perplexity do not always agree on model ranking; (iii) when perplexity and benchmark score disagree, the latter better reflects factuality in open-ended generation, as measured by human annotators. We make our data and code publicly available in https://github.com/AI21Labs/factor.
format Preprint
id arxiv_https___arxiv_org_abs_2307_06908
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generating Benchmarks for Factuality Evaluation of Language Models
Muhlgay, Dor
Ram, Ori
Magar, Inbal
Levine, Yoav
Ratner, Nir
Belinkov, Yonatan
Abend, Omri
Leyton-Brown, Kevin
Shashua, Amnon
Shoham, Yoav
Computation and Language
Artificial Intelligence
Before deploying a language model (LM) within a given domain, it is important to measure its tendency to generate factually incorrect information in that domain. Existing methods for factuality evaluation of LLM generation focus on facts sampled from the LM itself, and thus do not control the set of evaluated facts and might under-represent domain specific or rare facts. We propose FACTOR: Factual Assessment via Corpus TransfORmation, a scalable approach for evaluating LM factuality. FACTOR automatically transforms a factual corpus of interest into a benchmark evaluating an LM's propensity to generate true facts from the corpus vs. similar but incorrect statements. We use our framework to create three benchmarks: Wiki-FACTOR, News-FACTOR and Expert-FACTOR. We show that: (i) our benchmark scores increase with model size and improve when the LM is augmented with retrieval; (ii) benchmark score and perplexity do not always agree on model ranking; (iii) when perplexity and benchmark score disagree, the latter better reflects factuality in open-ended generation, as measured by human annotators. We make our data and code publicly available in https://github.com/AI21Labs/factor.
title Generating Benchmarks for Factuality Evaluation of Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2307.06908