Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lotfi, Sanae, Kuang, Yilun, Amos, Brandon, Goldblum, Micah, Finzi, Marc, Wilson, Andrew Gordon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917733458771968
author Lotfi, Sanae
Kuang, Yilun
Amos, Brandon
Goldblum, Micah
Finzi, Marc
Wilson, Andrew Gordon
author_facet Lotfi, Sanae
Kuang, Yilun
Amos, Brandon
Goldblum, Micah
Finzi, Marc
Wilson, Andrew Gordon
contents Large language models (LLMs) with billions of parameters excel at predicting the next token in a sequence. Recent work computes non-vacuous compression-based generalization bounds for LLMs, but these bounds are vacuous for large models at the billion-parameter scale. Moreover, these bounds are obtained through restrictive compression techniques, bounding compressed models that generate low-quality text. Additionally, the tightness of these existing bounds depends on the number of IID documents in a training set rather than the much larger number of non-IID constituent tokens, leaving untapped potential for tighter bounds. In this work, we instead use properties of martingales to derive generalization bounds that benefit from the vast number of tokens in LLM training sets. Since a dataset contains far more tokens than documents, our generalization bounds not only tolerate but actually benefit from far less restrictive compression schemes. With Monarch matrices, Kronecker factorizations, and post-training quantization, we achieve non-vacuous generalization bounds for LLMs as large as LLaMA2-70B. Unlike previous approaches, our work achieves the first non-vacuous bounds for models that are deployed in practice and generate high-quality text.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18158
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
Lotfi, Sanae
Kuang, Yilun
Amos, Brandon
Goldblum, Micah
Finzi, Marc
Wilson, Andrew Gordon
Machine Learning
Large language models (LLMs) with billions of parameters excel at predicting the next token in a sequence. Recent work computes non-vacuous compression-based generalization bounds for LLMs, but these bounds are vacuous for large models at the billion-parameter scale. Moreover, these bounds are obtained through restrictive compression techniques, bounding compressed models that generate low-quality text. Additionally, the tightness of these existing bounds depends on the number of IID documents in a training set rather than the much larger number of non-IID constituent tokens, leaving untapped potential for tighter bounds. In this work, we instead use properties of martingales to derive generalization bounds that benefit from the vast number of tokens in LLM training sets. Since a dataset contains far more tokens than documents, our generalization bounds not only tolerate but actually benefit from far less restrictive compression schemes. With Monarch matrices, Kronecker factorizations, and post-training quantization, we achieve non-vacuous generalization bounds for LLMs as large as LLaMA2-70B. Unlike previous approaches, our work achieves the first non-vacuous bounds for models that are deployed in practice and generate high-quality text.
title Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
topic Machine Learning
url https://arxiv.org/abs/2407.18158