Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Scott, Hansen, Dutch, Li, Jerry
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915987214827520
author Geng, Scott
Hansen, Dutch
Li, Jerry
author_facet Geng, Scott
Hansen, Dutch
Li, Jerry
contents Weak-to-strong generalization is a phenomenon in post-training whereby a strong student model, when finetuned solely with feedback from a weaker teacher, can not only surpass the teacher, but can improve upon its own capabilities. Recent work of Burns et al. (2023) demonstrated that this can occur in the setting of frontier language models, and subsequently there has been a flurry of both empirical work trying to exploit this phenomenon, as well as theoretical work attempting to understand it. In this work, we demonstrate that weak-to-strong generalization occurs in standard linear logistic regression, under mild distributional assumptions on the data. In fact, we show that this happens for most student-teacher pairs, suggesting that weak-to-strong generalization is in fact \emph{almost inevitable}, even in this basic setting. Notably, our setting does not require the student to be more expressive or have more model capacity in any way compared to the teacher, which runs contrary to the prevailing theoretical belief that a mismatch in model capacity is a central mechanism to weak-to-strong generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05742
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)
Geng, Scott
Hansen, Dutch
Li, Jerry
Machine Learning
Weak-to-strong generalization is a phenomenon in post-training whereby a strong student model, when finetuned solely with feedback from a weaker teacher, can not only surpass the teacher, but can improve upon its own capabilities. Recent work of Burns et al. (2023) demonstrated that this can occur in the setting of frontier language models, and subsequently there has been a flurry of both empirical work trying to exploit this phenomenon, as well as theoretical work attempting to understand it. In this work, we demonstrate that weak-to-strong generalization occurs in standard linear logistic regression, under mild distributional assumptions on the data. In fact, we show that this happens for most student-teacher pairs, suggesting that weak-to-strong generalization is in fact \emph{almost inevitable}, even in this basic setting. Notably, our setting does not require the student to be more expressive or have more model capacity in any way compared to the teacher, which runs contrary to the prevailing theoretical belief that a mismatch in model capacity is a central mechanism to weak-to-strong generalization.
title Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)
topic Machine Learning
url https://arxiv.org/abs/2605.05742