Saved in:
| Main Authors: | Beck, Alon, Sinai, Yohai Bar, Levi, Noam |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.12039 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grokking at the Edge of Linear Separability
by: Beck, Alon, et al.
Published: (2024)
by: Beck, Alon, et al.
Published: (2024)
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
by: Levi, Noam, et al.
Published: (2023)
by: Levi, Noam, et al.
Published: (2023)
Hidden Markov modeling of single particle diffusion with stochastic tethering
by: Federbush, Amit, et al.
Published: (2023)
by: Federbush, Amit, et al.
Published: (2023)
A Simple Model of Inference Scaling Laws
by: Levi, Noam
Published: (2024)
by: Levi, Noam
Published: (2024)
The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels
by: Slutzky, Yonatan, et al.
Published: (2024)
by: Slutzky, Yonatan, et al.
Published: (2024)
Learning Shrinks the Hard Tail: Training-Dependent Inference Scaling in a Solvable Linear Model
by: Levi, Noam
Published: (2026)
by: Levi, Noam
Published: (2026)
Machine Learning the Entropy to Estimate Free Energy Differences without Sampling Transitions
by: Ben-Shimon, Yamin, et al.
Published: (2025)
by: Ben-Shimon, Yamin, et al.
Published: (2025)
Mirror, Mirror of the Flow: How Does Regularization Shape Implicit Bias?
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Consistency Regularization for Domain Generalization with Logit Attribution Matching
by: Gao, Han, et al.
Published: (2023)
by: Gao, Han, et al.
Published: (2023)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
by: Bach, Thong, et al.
Published: (2026)
by: Bach, Thong, et al.
Published: (2026)
Classifying Overlapping Gaussian Mixtures in High Dimensions: From Optimal Classifiers to Neural Nets
by: Cohen, Khen, et al.
Published: (2024)
by: Cohen, Khen, et al.
Published: (2024)
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
The Underlying Scaling Laws and Universal Statistical Structure of Complex Datasets
by: Levi, Noam, et al.
Published: (2023)
by: Levi, Noam, et al.
Published: (2023)
Pretraining Scaling Laws for Generative Evaluations of Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Decoupled Weight Decay for Any $p$ Norm
by: Outmezguine, Nadav Joseph, et al.
Published: (2024)
by: Outmezguine, Nadav Joseph, et al.
Published: (2024)
Random Initialization of Gated Sparse Adapters
by: Retault, Vi, et al.
Published: (2025)
by: Retault, Vi, et al.
Published: (2025)
GRIT: Graph-Regularized Logit Refinement for Zero-shot Cell Type Annotation
by: Hu, Tianxiang, et al.
Published: (2025)
by: Hu, Tianxiang, et al.
Published: (2025)
Earthquake magnitudes depend on seismic history, as revealed by a neural network analysis
by: Berman, Neri, et al.
Published: (2024)
by: Berman, Neri, et al.
Published: (2024)
Optimal Implicit Bias in Linear Regression
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
Estimating Implicit Regularization in Deep Learning
by: Rudoler, Joseph H., et al.
Published: (2026)
by: Rudoler, Joseph H., et al.
Published: (2026)
On the Implicit Bias of Adam
by: Cattaneo, Matias D., et al.
Published: (2023)
by: Cattaneo, Matias D., et al.
Published: (2023)
Ascent Fails to Forget
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
by: Mavrothalassitis, Ioannis, et al.
Published: (2025)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Implicit Bias in Deep Linear Discriminant Analysis
by: Li, Jiawen
Published: (2026)
by: Li, Jiawen
Published: (2026)
The Price of Implicit Bias in Adversarially Robust Generalization
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
by: Jacobs, Tom, et al.
Published: (2025)
by: Jacobs, Tom, et al.
Published: (2025)
Can Implicit Bias Imply Adversarial Robustness?
by: Min, Hancheng, et al.
Published: (2024)
by: Min, Hancheng, et al.
Published: (2024)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
Logits-Based Finetuning
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
Measuring Sharpness in Grokking
by: Miller, Jack, et al.
Published: (2024)
by: Miller, Jack, et al.
Published: (2024)
Implicit Bias of the JKO Scheme
by: Halmos, Peter, et al.
Published: (2025)
by: Halmos, Peter, et al.
Published: (2025)
Scaled Supervision is an Implicit Lipschitz Regularizer
by: Ouyang, Zhongyu, et al.
Published: (2025)
by: Ouyang, Zhongyu, et al.
Published: (2025)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
Implicit Regularization of the Deep Inverse Prior Trained with Inertia
by: Buskulic, Nathan, et al.
Published: (2025)
by: Buskulic, Nathan, et al.
Published: (2025)
Why is Your Language Model a Poor Implicit Reward Model?
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning
by: Tang, Yuwei, et al.
Published: (2024)
by: Tang, Yuwei, et al.
Published: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Similar Items
-
Grokking at the Edge of Linear Separability
by: Beck, Alon, et al.
Published: (2024) -
Grokking in Linear Estimators -- A Solvable Model that Groks without Understanding
by: Levi, Noam, et al.
Published: (2023) -
Hidden Markov modeling of single particle diffusion with stochastic tethering
by: Federbush, Amit, et al.
Published: (2023) -
A Simple Model of Inference Scaling Laws
by: Levi, Noam
Published: (2024) -
The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels
by: Slutzky, Yonatan, et al.
Published: (2024)