On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Moniri, Behrad, Hassani, Hamed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026)
by: Moniri, Behrad, et al.
Published: (2026)
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023)
by: Moniri, Behrad, et al.
Published: (2023)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
by: Zhang, Thomas T., et al.
Published: (2025)
by: Zhang, Thomas T., et al.
Published: (2025)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
by: Kiyani, Shayan, et al.
Published: (2026)
by: Kiyani, Shayan, et al.
Published: (2026)
Theoretical Analysis of Weak-to-Strong Generalization
by: Lang, Hunter, et al.
Published: (2024)
by: Lang, Hunter, et al.
Published: (2024)
Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models
by: Noorani, Sima, et al.
Published: (2025)
by: Noorani, Sima, et al.
Published: (2025)
On the Emergence of Weak-to-Strong Generalization: A Bias-Variance Perspective
by: Xu, Gengze, et al.
Published: (2025)
by: Xu, Gengze, et al.
Published: (2025)
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
Generalization Properties of Adversarial Training for $\ell_0$-Bounded Adversarial Attacks
by: Delgosha, Payam, et al.
Published: (2024)
by: Delgosha, Payam, et al.
Published: (2024)
The curse of overparametrization in adversarial training: Precise analysis of robust generalization for random features regression
by: Hassani, Hamed, et al.
Published: (2022)
by: Hassani, Hamed, et al.
Published: (2022)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
Watermark Smoothing Attacks against Language Models
by: Chang, Hongyan, et al.
Published: (2024)
by: Chang, Hongyan, et al.
Published: (2024)
On Weak-to-Strong Generalization and f-Divergence
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
Explicitly Encoding Structural Symmetry is Key to Length Generalization in Arithmetic Tasks
by: Sabbaghi, Mahdi, et al.
Published: (2024)
by: Sabbaghi, Mahdi, et al.
Published: (2024)
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
by: Yao, Wei, et al.
Published: (2025)
by: Yao, Wei, et al.
Published: (2025)
On the Blessing of Pre-training in Weak-to-Strong Generalization
by: Yao, Wei, et al.
Published: (2026)
by: Yao, Wei, et al.
Published: (2026)
Conformal Prediction with Learned Features
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
Quantifying the Gain in Weak-to-Strong Generalization
by: Charikar, Moses, et al.
Published: (2024)
by: Charikar, Moses, et al.
Published: (2024)
Length Optimization in Conformal Prediction
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
Provable Weak-to-Strong Generalization via Benign Overfitting
by: Wu, David X., et al.
Published: (2024)
by: Wu, David X., et al.
Published: (2024)
Weak-to-Strong Generalization is Nearly Inevitable (in Linear Models)
by: Geng, Scott, et al.
Published: (2026)
by: Geng, Scott, et al.
Published: (2026)
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
by: Joshi, Sunay, et al.
Published: (2025)
by: Joshi, Sunay, et al.
Published: (2025)
Provable tradeoffs in adversarially robust classification
by: Dobriban, Edgar, et al.
Published: (2020)
by: Dobriban, Edgar, et al.
Published: (2020)
InfoSFT: Learn More and Forget Less with Information-Aware Token Weighting
by: Sabbaghi, Mahdi, et al.
Published: (2026)
by: Sabbaghi, Mahdi, et al.
Published: (2026)
Multi-Round Human-AI Collaboration with User-Specified Requirements
by: Noorani, Sima, et al.
Published: (2026)
by: Noorani, Sima, et al.
Published: (2026)
Robust Policy Optimization to Prevent Catastrophic Forgetting
by: Sabbaghi, Mahdi, et al.
Published: (2026)
by: Sabbaghi, Mahdi, et al.
Published: (2026)
Optimal Neural Compressors for the Rate-Distortion-Perception Tradeoff
by: Lei, Eric, et al.
Published: (2025)
by: Lei, Eric, et al.
Published: (2025)
Neural Estimation of the Rate-Distortion Function With Applications to Operational Source Coding
by: Lei, Eric, et al.
Published: (2022)
by: Lei, Eric, et al.
Published: (2022)
Weak-to-Strong Generalization under Distribution Shifts
by: Jeon, Myeongho, et al.
Published: (2025)
by: Jeon, Myeongho, et al.
Published: (2025)
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
by: Ye, Ruimeng, et al.
Published: (2025)
by: Ye, Ruimeng, et al.
Published: (2025)
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
by: Pawelczyk, Martin, et al.
Published: (2024)
by: Pawelczyk, Martin, et al.
Published: (2024)
Weak-to-Strong Generalization Through the Data-Centric Lens
by: Shin, Changho, et al.
Published: (2024)
by: Shin, Changho, et al.
Published: (2024)
Robust Decision Making with Partially Calibrated Forecasts
by: Kiyani, Shayan, et al.
Published: (2025)
by: Kiyani, Shayan, et al.
Published: (2025)
Risk-Controlled Post-Processing of Decision Policies
by: Joshi, Sunay, et al.
Published: (2026)
by: Joshi, Sunay, et al.
Published: (2026)
Compression of Structured Data with Autoencoders: Provable Benefit of Nonlinearities and Depth
by: Kögler, Kevin, et al.
Published: (2024)
by: Kögler, Kevin, et al.
Published: (2024)
Similar Items
-
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026) -
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024) -
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024) -
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024) -
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023)