Salvato in:
| Autori principali: | Medvedev, Marko, Lyu, Kaifeng, Yu, Dingli, Arora, Sanjeev, Li, Zhiyuan, Srebro, Nathan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.02877 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Shift is Good: Mismatched Data Mixing Improves Test Performance
di: Medvedev, Marko, et al.
Pubblicazione: (2025)
di: Medvedev, Marko, et al.
Pubblicazione: (2025)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
di: Medvedev, Marko, et al.
Pubblicazione: (2024)
di: Medvedev, Marko, et al.
Pubblicazione: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
di: Lyu, Kaifeng, et al.
Pubblicazione: (2024)
di: Lyu, Kaifeng, et al.
Pubblicazione: (2024)
Positive Distribution Shift as a Framework for Understanding Tractable Learning
di: Medvedev, Marko, et al.
Pubblicazione: (2026)
di: Medvedev, Marko, et al.
Pubblicazione: (2026)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
di: Malladi, Sadhika, et al.
Pubblicazione: (2022)
di: Malladi, Sadhika, et al.
Pubblicazione: (2022)
Recursive Models for Long-Horizon Reasoning
di: Yang, Chenxiao, et al.
Pubblicazione: (2026)
di: Yang, Chenxiao, et al.
Pubblicazione: (2026)
AI-Assisted Generation of Difficult Math Questions
di: Shah, Vedant, et al.
Pubblicazione: (2024)
di: Shah, Vedant, et al.
Pubblicazione: (2024)
A Quadratic Synchronization Rule for Distributed Deep Learning
di: Gu, Xinran, et al.
Pubblicazione: (2023)
di: Gu, Xinran, et al.
Pubblicazione: (2023)
Can Models Learn Skill Composition from Examples?
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
di: Zhao, Haoyu, et al.
Pubblicazione: (2024)
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
di: Oh, Junsoo, et al.
Pubblicazione: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
di: Park, Simon, et al.
Pubblicazione: (2025)
di: Park, Simon, et al.
Pubblicazione: (2025)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
di: Harel, Itamar, et al.
Pubblicazione: (2024)
di: Harel, Itamar, et al.
Pubblicazione: (2024)
Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking
di: Lyu, Kaifeng, et al.
Pubblicazione: (2023)
di: Lyu, Kaifeng, et al.
Pubblicazione: (2023)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
di: Joshi, Nirmit, et al.
Pubblicazione: (2023)
di: Joshi, Nirmit, et al.
Pubblicazione: (2023)
PENCIL: Long Thoughts with Short Memory
di: Yang, Chenxiao, et al.
Pubblicazione: (2025)
di: Yang, Chenxiao, et al.
Pubblicazione: (2025)
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
di: Zhu, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaohan, et al.
Pubblicazione: (2025)
Provable Weak-to-Strong Generalization via Benign Overfitting
di: Wu, David X., et al.
Pubblicazione: (2024)
di: Wu, David X., et al.
Pubblicazione: (2024)
Provable unlearning in topic modeling and downstream tasks
di: Wei, Stanley, et al.
Pubblicazione: (2024)
di: Wei, Stanley, et al.
Pubblicazione: (2024)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
di: Buzaglo, Gon, et al.
Pubblicazione: (2024)
di: Buzaglo, Gon, et al.
Pubblicazione: (2024)
Overfitting and Generalizing with (PAC) Bayesian Prediction in Noisy Binary Classification
di: Zhu, Xiaohan, et al.
Pubblicazione: (2026)
di: Zhu, Xiaohan, et al.
Pubblicazione: (2026)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
di: Li, Binghui, et al.
Pubblicazione: (2024)
di: Li, Binghui, et al.
Pubblicazione: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2024)
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2024)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
di: Wu, Diyuan, et al.
Pubblicazione: (2026)
di: Wu, Diyuan, et al.
Pubblicazione: (2026)
Depth Separation in Norm-Bounded Infinite-Width Neural Networks
di: Parkinson, Suzanna, et al.
Pubblicazione: (2024)
di: Parkinson, Suzanna, et al.
Pubblicazione: (2024)
Strong and Weak Random Walks on Signed Networks
di: Babul, Shazia'Ayn, et al.
Pubblicazione: (2024)
di: Babul, Shazia'Ayn, et al.
Pubblicazione: (2024)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
di: Joshi, Nirmit, et al.
Pubblicazione: (2024)
di: Joshi, Nirmit, et al.
Pubblicazione: (2024)
Tight Bounds on the Binomial CDF, and the Minimum of i.i.d Binomials, in terms of KL-Divergence
di: Zhu, Xiaohan, et al.
Pubblicazione: (2025)
di: Zhu, Xiaohan, et al.
Pubblicazione: (2025)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
di: Harel, Itamar, et al.
Pubblicazione: (2025)
di: Harel, Itamar, et al.
Pubblicazione: (2025)
Research Program: Theory of Learning in Dynamical Systems
di: Hazan, Elad, et al.
Pubblicazione: (2025)
di: Hazan, Elad, et al.
Pubblicazione: (2025)
The Marginal Value of Momentum for Small Learning Rate SGD
di: Wang, Runzhe, et al.
Pubblicazione: (2023)
di: Wang, Runzhe, et al.
Pubblicazione: (2023)
A Theory of Learning with Autoregressive Chain of Thought
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
di: Zhou, Lijia, et al.
Pubblicazione: (2023)
di: Zhou, Lijia, et al.
Pubblicazione: (2023)
On the Hardness of Learning Regular Expressions
di: Attias, Idan, et al.
Pubblicazione: (2025)
di: Attias, Idan, et al.
Pubblicazione: (2025)
Learning single-index models via harmonic decomposition
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
di: Joshi, Nirmit, et al.
Pubblicazione: (2025)
Provable Benefits of Sinusoidal Activation for Modular Addition
di: Huang, Tianlong, et al.
Pubblicazione: (2025)
di: Huang, Tianlong, et al.
Pubblicazione: (2025)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
di: Awano, Ryoya, et al.
Pubblicazione: (2026)
di: Awano, Ryoya, et al.
Pubblicazione: (2026)
Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer Manifold
di: Li, Xinghan, et al.
Pubblicazione: (2025)
di: Li, Xinghan, et al.
Pubblicazione: (2025)
Mixture of Weak & Strong Experts on Graphs
di: Zeng, Hanqing, et al.
Pubblicazione: (2023)
di: Zeng, Hanqing, et al.
Pubblicazione: (2023)
On the Impossibility of Retrain Equivalence in Machine Unlearning
di: Yu, Jiatong, et al.
Pubblicazione: (2025)
di: Yu, Jiatong, et al.
Pubblicazione: (2025)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
di: Park, Simon, et al.
Pubblicazione: (2025)
di: Park, Simon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Shift is Good: Mismatched Data Mixing Improves Test Performance
di: Medvedev, Marko, et al.
Pubblicazione: (2025) -
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
di: Medvedev, Marko, et al.
Pubblicazione: (2024) -
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
di: Lyu, Kaifeng, et al.
Pubblicazione: (2024) -
Positive Distribution Shift as a Framework for Understanding Tractable Learning
di: Medvedev, Marko, et al.
Pubblicazione: (2026) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
di: Malladi, Sadhika, et al.
Pubblicazione: (2022)