Paloma: A Benchmark for Evaluating Language Model Fit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Magnusson, Ian, Bhagia, Akshita, Hofmann, Valentin, Soldaini, Luca, Jha, Ananya Harsh, Tafjord, Oyvind, Schwenk, Dustin, Walsh, Evan Pete, Elazar, Yanai, Lo, Kyle, Groeneveld, Dirk, Beltagy, Iz, Hajishirzi, Hannaneh, Smith, Noah A., Richardson, Kyle, Dodge, Jesse
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916511539527680
author Magnusson, Ian
Bhagia, Akshita
Hofmann, Valentin
Soldaini, Luca
Jha, Ananya Harsh
Tafjord, Oyvind
Schwenk, Dustin
Walsh, Evan Pete
Elazar, Yanai
Lo, Kyle
Groeneveld, Dirk
Beltagy, Iz
Hajishirzi, Hannaneh
Smith, Noah A.
Richardson, Kyle
Dodge, Jesse
author_facet Magnusson, Ian
Bhagia, Akshita
Hofmann, Valentin
Soldaini, Luca
Jha, Ananya Harsh
Tafjord, Oyvind
Schwenk, Dustin
Walsh, Evan Pete
Elazar, Yanai
Lo, Kyle
Groeneveld, Dirk
Beltagy, Iz
Hajishirzi, Hannaneh
Smith, Noah A.
Richardson, Kyle
Dodge, Jesse
contents Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains--varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10523
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Paloma: A Benchmark for Evaluating Language Model Fit
Magnusson, Ian
Bhagia, Akshita
Hofmann, Valentin
Soldaini, Luca
Jha, Ananya Harsh
Tafjord, Oyvind
Schwenk, Dustin
Walsh, Evan Pete
Elazar, Yanai
Lo, Kyle
Groeneveld, Dirk
Beltagy, Iz
Hajishirzi, Hannaneh
Smith, Noah A.
Richardson, Kyle
Dodge, Jesse
Computation and Language
Artificial Intelligence
Machine Learning
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains--varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. We include two new datasets of the top 100 subreddits (e.g., r/depression on Reddit) and programming languages (e.g., Java on GitHub), both sources common in contemporary LMs. With our benchmark, we release 6 baseline 1B LMs carefully controlled to provide fair comparisons about which pretraining corpus is best and code for others to apply those controls to their own experiments. Our case studies demonstrate how the fine-grained results from Paloma surface findings such as that models pretrained without data beyond Common Crawl exhibit anomalous gaps in LM fit to many domains or that loss is dominated by the most frequently occurring strings in the vocabulary.
title Paloma: A Benchmark for Evaluating Language Model Fit
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.10523