The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Awano, Ryoya, Suzuki, Taiji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
von: Ye, Ruimeng, et al.
Veröffentlicht: (2025)
von: Ye, Ruimeng, et al.
Veröffentlicht: (2025)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2026)
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
von: Deng, Wei
Veröffentlicht: (2026)
von: Deng, Wei
Veröffentlicht: (2026)
Deep Two-Way Matrix Reordering for Relational Data Analysis
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
von: Watanabe, Chihiro, et al.
Veröffentlicht: (2021)
Test time training enhances in-context learning of nonlinear functions
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
von: Takakura, Shokichi, et al.
Veröffentlicht: (2023)
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
von: Moniri, Behrad, et al.
Veröffentlicht: (2025)
von: Moniri, Behrad, et al.
Veröffentlicht: (2025)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
von: Medvedev, Marko, et al.
Veröffentlicht: (2025)
von: Medvedev, Marko, et al.
Veröffentlicht: (2025)
Eliciting Latent Knowledge from Quirky Language Models
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
von: Mallen, Alex, et al.
Veröffentlicht: (2023)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
von: Higuchi, Rei, et al.
Veröffentlicht: (2025)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
von: Kim, Juno, et al.
Veröffentlicht: (2024)
von: Kim, Juno, et al.
Veröffentlicht: (2024)
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
von: Oh, Junsoo, et al.
Veröffentlicht: (2025)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2026)
von: Lascu, Razvan-Andrei, et al.
Veröffentlicht: (2026)
Approximately Unimodal Likelihood Models for Ordinal Regression
von: Yamasaki, Ryoya
Veröffentlicht: (2025)
von: Yamasaki, Ryoya
Veröffentlicht: (2025)
High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
von: Fu, Guoji, et al.
Veröffentlicht: (2026)
von: Fu, Guoji, et al.
Veröffentlicht: (2026)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
von: Wu, Diyuan, et al.
Veröffentlicht: (2026)
von: Wu, Diyuan, et al.
Veröffentlicht: (2026)
How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
von: Higuchi, Rei, et al.
Veröffentlicht: (2026)
Eliciting Latent Predictions from Transformers with the Tuned Lens
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
von: Belrose, Nora, et al.
Veröffentlicht: (2023)
On Weak-to-Strong Generalization and f-Divergence
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
von: Yamamoto, Naoya, et al.
Veröffentlicht: (2025)
On the Role of Label Noise in the Feature Learning Process
von: Han, Andi, et al.
Veröffentlicht: (2025)
von: Han, Andi, et al.
Veröffentlicht: (2025)
Eliciting Secret Knowledge from Language Models
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
von: Yao, Wei, et al.
Veröffentlicht: (2025)
von: Yao, Wei, et al.
Veröffentlicht: (2025)
On the Blessing of Pre-training in Weak-to-Strong Generalization
von: Yao, Wei, et al.
Veröffentlicht: (2026)
von: Yao, Wei, et al.
Veröffentlicht: (2026)
Zero-Flow Encoders
von: Wang, Yakun, et al.
Veröffentlicht: (2026)
von: Wang, Yakun, et al.
Veröffentlicht: (2026)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
von: Kim, Juno, et al.
Veröffentlicht: (2024) -
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
von: Wakayama, Tomoya, et al.
Veröffentlicht: (2025) -
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
von: Nishikawa, Naoki, et al.
Veröffentlicht: (2025) -
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
von: Oh, Junsoo, et al.
Veröffentlicht: (2025) -
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
von: Ye, Ruimeng, et al.
Veröffentlicht: (2025)