Saved in:
Bibliographic Details
Main Authors: Awano, Ryoya, Suzuki, Taiji
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.12908
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916008531329024
author Awano, Ryoya
Suzuki, Taiji
author_facet Awano, Ryoya
Suzuki, Taiji
contents Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $κ$. We prove that the strong model efficiently learns task $κ$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12908
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
Awano, Ryoya
Suzuki, Taiji
Machine Learning
Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $κ$. We prove that the strong model efficiently learns task $κ$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.
title The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
topic Machine Learning
url https://arxiv.org/abs/2605.12908