Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miranda, Brando, Lee, Alycia, Sundar, Sudharsan, Casasola, Allison, Schaeffer, Rylan, Obbad, Elyas, Koyejo, Sanmi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911035480342528
author Miranda, Brando
Lee, Alycia
Sundar, Sudharsan
Casasola, Allison
Schaeffer, Rylan
Obbad, Elyas
Koyejo, Sanmi
author_facet Miranda, Brando
Lee, Alycia
Sundar, Sudharsan
Casasola, Allison
Schaeffer, Rylan
Obbad, Elyas
Koyejo, Sanmi
contents Current trends in pre-training Large Language Models (LLMs) primarily focus on the scaling of model and dataset size. While the quality of pre-training data is considered an important factor for training powerful LLMs, it remains a nebulous concept that has not been rigorously characterized. To this end, we propose a formalization of one key aspect of data quality -- measuring the variability of natural language data -- specifically via a measure we call the diversity coefficient. Our empirical analysis shows that the proposed diversity coefficient aligns with the intuitive properties of diversity and variability, e.g., it increases as the number of latent concepts increases. Then, we measure the diversity coefficient of publicly available pre-training datasets and demonstrate that their formal diversity is high compared to theoretical lower and upper bounds. Finally, we conduct a comprehensive set of controlled interventional experiments with GPT-2 and LLaMAv2 that demonstrate the diversity coefficient of pre-training data characterizes useful aspects of downstream model evaluation performance -- totaling 44 models of various sizes (51M to 7B parameters). We conclude that our formal notion of diversity is an important aspect of data quality that captures variability and causally leads to improved evaluation performance.
format Preprint
id arxiv_https___arxiv_org_abs_2306_13840
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
Miranda, Brando
Lee, Alycia
Sundar, Sudharsan
Casasola, Allison
Schaeffer, Rylan
Obbad, Elyas
Koyejo, Sanmi
Computation and Language
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
Current trends in pre-training Large Language Models (LLMs) primarily focus on the scaling of model and dataset size. While the quality of pre-training data is considered an important factor for training powerful LLMs, it remains a nebulous concept that has not been rigorously characterized. To this end, we propose a formalization of one key aspect of data quality -- measuring the variability of natural language data -- specifically via a measure we call the diversity coefficient. Our empirical analysis shows that the proposed diversity coefficient aligns with the intuitive properties of diversity and variability, e.g., it increases as the number of latent concepts increases. Then, we measure the diversity coefficient of publicly available pre-training datasets and demonstrate that their formal diversity is high compared to theoretical lower and upper bounds. Finally, we conduct a comprehensive set of controlled interventional experiments with GPT-2 and LLaMAv2 that demonstrate the diversity coefficient of pre-training data characterizes useful aspects of downstream model evaluation performance -- totaling 44 models of various sizes (51M to 7B parameters). We conclude that our formal notion of diversity is an important aspect of data quality that captures variability and causally leads to improved evaluation performance.
title Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
topic Computation and Language
Artificial Intelligence
Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2306.13840