Positive Distribution Shift as a Framework for Understanding Tractable Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Medvedev, Marko, Attias, Idan, Cornacchia, Elisabetta, Misiakiewicz, Theodor, Vardi, Gal, Srebro, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
by: Joshi, Nirmit, et al.
Published: (2024)
by: Joshi, Nirmit, et al.
Published: (2024)
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
Learning single-index models via harmonic decomposition
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
Shift is Good: Mismatched Data Mixing Improves Test Performance
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023)
by: Zhou, Lijia, et al.
Published: (2023)
Learning to Think from Multiple Thinkers
by: Joshi, Nirmit, et al.
Published: (2026)
by: Joshi, Nirmit, et al.
Published: (2026)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
A non-asymptotic theory of Kernel Ridge Regression: deterministic equivalents, test error, and GCV estimator
by: Misiakiewicz, Theodor, et al.
Published: (2024)
by: Misiakiewicz, Theodor, et al.
Published: (2024)
Statistical-Computational Trade-offs in Learning Multi-Index Models via Harmonic Analysis
by: Latourelle-Vigeant, Hugo, et al.
Published: (2026)
by: Latourelle-Vigeant, Hugo, et al.
Published: (2026)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
A Mathematical Model for Curriculum Learning for Parities
by: Cornacchia, Elisabetta, et al.
Published: (2023)
by: Cornacchia, Elisabetta, et al.
Published: (2023)
Learning with Shallow Neural Networks on Cluster-Structured Features
by: Cornacchia, Elisabetta, et al.
Published: (2026)
by: Cornacchia, Elisabetta, et al.
Published: (2026)
Capacity-Constrained Online Learning with Delays: Scheduling Frameworks and Regret Trade-offs
by: Ryabchenko, Alexander, et al.
Published: (2025)
by: Ryabchenko, Alexander, et al.
Published: (2025)
Adversarially Robust PAC Learnability of Real-Valued Functions
by: Attias, Idan, et al.
Published: (2022)
by: Attias, Idan, et al.
Published: (2022)
Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions
by: Cornacchia, Elisabetta, et al.
Published: (2025)
by: Cornacchia, Elisabetta, et al.
Published: (2025)
Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning
by: Attias, Idan, et al.
Published: (2026)
by: Attias, Idan, et al.
Published: (2026)
Dimension-free deterministic equivalents and scaling laws for random feature regression
by: Defilippis, Leonardo, et al.
Published: (2024)
by: Defilippis, Leonardo, et al.
Published: (2024)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
Asymptotics of Random Feature Regression Beyond the Linear Scaling Regime
by: Hu, Hong, et al.
Published: (2024)
by: Hu, Hong, et al.
Published: (2024)
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
by: Wu, Diyuan, et al.
Published: (2026)
by: Wu, Diyuan, et al.
Published: (2026)
The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently
by: Cornacchia, Elisabetta, et al.
Published: (2026)
by: Cornacchia, Elisabetta, et al.
Published: (2026)
Tradeoffs between Mistakes and ERM Oracle Calls in Online and Transductive Online Learning
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Sample Compression Scheme Reductions
by: Attias, Idan, et al.
Published: (2024)
by: Attias, Idan, et al.
Published: (2024)
A Characterization of Semi-Supervised Adversarially-Robust PAC Learnability
by: Attias, Idan, et al.
Published: (2022)
by: Attias, Idan, et al.
Published: (2022)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
Transformers are almost optimal metalearners for linear classification
by: Magen, Roey, et al.
Published: (2025)
by: Magen, Roey, et al.
Published: (2025)
Learning-Augmented Algorithms for Boolean Satisfiability
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Causal Bandits: The Pareto Optimal Frontier of Adaptivity, a Reduction to Linear Bandits, and Limitations around Unknown Marginals
by: Liu, Ziyi, et al.
Published: (2024)
by: Liu, Ziyi, et al.
Published: (2024)
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
by: Zhu, Xiaohan, et al.
Published: (2025)
by: Zhu, Xiaohan, et al.
Published: (2025)
Sequential Probability Assignment with Contexts: Minimax Regret, Contextual Shtarkov Sums, and Contextual Normalized Maximum Likelihood
by: Liu, Ziyi, et al.
Published: (2024)
by: Liu, Ziyi, et al.
Published: (2024)
An Optimized Franz-Parisi Criterion and its Equivalence with SQ Lower Bounds
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
Learning High-Degree Parities: The Crucial Role of the Initialization
by: Abbe, Emmanuel, et al.
Published: (2024)
by: Abbe, Emmanuel, et al.
Published: (2024)
Optimal Learners for Realizable Regression: PAC Learning and Online Learning
by: Attias, Idan, et al.
Published: (2023)
by: Attias, Idan, et al.
Published: (2023)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
Similar Items
-
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024) -
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025) -
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
by: Joshi, Nirmit, et al.
Published: (2024) -
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025) -
Learning single-index models via harmonic decomposition
by: Joshi, Nirmit, et al.
Published: (2025)