Quantifying the Gain in Weak-to-Strong Generalization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Charikar, Moses, Pabbaraju, Chirag, Shiragur, Kirankumar
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916449248870400
author Charikar, Moses
Pabbaraju, Chirag
Shiragur, Kirankumar
author_facet Charikar, Moses
Pabbaraju, Chirag
Shiragur, Kirankumar
contents Recent advances in large language models have shown capabilities that are extraordinary and near-superhuman. These models operate with such complexity that reliably evaluating and aligning them proves challenging for humans. This leads to the natural question: can guidance from weak models (like humans) adequately direct the capabilities of strong models? In a recent and somewhat surprising work, Burns et al. (2023) empirically demonstrated that when strong models (like GPT-4) are finetuned using labels generated by weak supervisors (like GPT-2), the strong models outperform their weaker counterparts -- a phenomenon they term weak-to-strong generalization. In this work, we present a theoretical framework for understanding weak-to-strong generalization. Specifically, we show that the improvement in performance achieved by strong models over their weaker counterparts is quantified by the misfit error incurred by the strong model on labels generated by the weaker model. Our theory reveals several curious algorithmic insights. For instance, we can predict the amount by which the strong model will improve over the weak model, and also choose among different weak models to train the strong model, based on its misfit error. We validate our theoretical findings through various empirical assessments.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15116
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Quantifying the Gain in Weak-to-Strong Generalization
Charikar, Moses
Pabbaraju, Chirag
Shiragur, Kirankumar
Machine Learning
Artificial Intelligence
Recent advances in large language models have shown capabilities that are extraordinary and near-superhuman. These models operate with such complexity that reliably evaluating and aligning them proves challenging for humans. This leads to the natural question: can guidance from weak models (like humans) adequately direct the capabilities of strong models? In a recent and somewhat surprising work, Burns et al. (2023) empirically demonstrated that when strong models (like GPT-4) are finetuned using labels generated by weak supervisors (like GPT-2), the strong models outperform their weaker counterparts -- a phenomenon they term weak-to-strong generalization. In this work, we present a theoretical framework for understanding weak-to-strong generalization. Specifically, we show that the improvement in performance achieved by strong models over their weaker counterparts is quantified by the misfit error incurred by the strong model on labels generated by the weaker model. Our theory reveals several curious algorithmic insights. For instance, we can predict the amount by which the strong model will improve over the weak model, and also choose among different weak models to train the strong model, based on its misfit error. We validate our theoretical findings through various empirical assessments.
title Quantifying the Gain in Weak-to-Strong Generalization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.15116