Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
Fuente:
arXiv
Guardado en:
| Autores principales: | Harel, Itamar, Wolanowsky, Yonathan, Vardi, Gal, Srebro, Nathan, Soudry, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Provable Tempered Overfitting of Minimal Nets and Typical Nets
por: Harel, Itamar, et al.
Publicado: (2024)
por: Harel, Itamar, et al.
Publicado: (2024)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
por: Buzaglo, Gon, et al.
Publicado: (2024)
por: Buzaglo, Gon, et al.
Publicado: (2024)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
por: Joshi, Nirmit, et al.
Publicado: (2023)
por: Joshi, Nirmit, et al.
Publicado: (2023)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
por: Medvedev, Marko, et al.
Publicado: (2024)
por: Medvedev, Marko, et al.
Publicado: (2024)
On the Hardness of Learning Regular Expressions
por: Attias, Idan, et al.
Publicado: (2025)
por: Attias, Idan, et al.
Publicado: (2025)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
por: Zhou, Lijia, et al.
Publicado: (2023)
por: Zhou, Lijia, et al.
Publicado: (2023)
Learning to Think from Multiple Thinkers
por: Joshi, Nirmit, et al.
Publicado: (2026)
por: Joshi, Nirmit, et al.
Publicado: (2026)
Positive Distribution Shift as a Framework for Understanding Tractable Learning
por: Medvedev, Marko, et al.
Publicado: (2026)
por: Medvedev, Marko, et al.
Publicado: (2026)
The Implicit Bias of Gradient Descent on Separable Data
por: Soudry, Daniel, et al.
Publicado: (2017)
por: Soudry, Daniel, et al.
Publicado: (2017)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
por: Frei, Spencer, et al.
Publicado: (2024)
por: Frei, Spencer, et al.
Publicado: (2024)
A Theory of Learning with Autoregressive Chain of Thought
por: Joshi, Nirmit, et al.
Publicado: (2025)
por: Joshi, Nirmit, et al.
Publicado: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
por: Evron, Itay, et al.
Publicado: (2025)
por: Evron, Itay, et al.
Publicado: (2025)
Transformers are almost optimal metalearners for linear classification
por: Magen, Roey, et al.
Publicado: (2025)
por: Magen, Roey, et al.
Publicado: (2025)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
por: Gronich, Eitan, et al.
Publicado: (2026)
por: Gronich, Eitan, et al.
Publicado: (2026)
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
por: Zhu, Xiaohan, et al.
Publicado: (2025)
por: Zhu, Xiaohan, et al.
Publicado: (2025)
Cross-Entropy Is All You Need To Invert the Data Generating Process
por: Reizinger, Patrik, et al.
Publicado: (2024)
por: Reizinger, Patrik, et al.
Publicado: (2024)
Overfitting and Generalizing with (PAC) Bayesian Prediction in Noisy Binary Classification
por: Zhu, Xiaohan, et al.
Publicado: (2026)
por: Zhu, Xiaohan, et al.
Publicado: (2026)
Research Program: Theory of Learning in Dynamical Systems
por: Hazan, Elad, et al.
Publicado: (2025)
por: Hazan, Elad, et al.
Publicado: (2025)
Accuracy is Not All You Need
por: Dutta, Abhinav, et al.
Publicado: (2024)
por: Dutta, Abhinav, et al.
Publicado: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
por: Tsilivis, Nikolaos, et al.
Publicado: (2024)
por: Tsilivis, Nikolaos, et al.
Publicado: (2024)
Implicit Regularization Towards Rank Minimization in ReLU Networks
por: Timor, Nadav, et al.
Publicado: (2022)
por: Timor, Nadav, et al.
Publicado: (2022)
To Grok Grokking: Provable Grokking in Ridge Regression
por: Xu, Mingyue, et al.
Publicado: (2026)
por: Xu, Mingyue, et al.
Publicado: (2026)
Realizable Learning is All You Need
por: Hopkins, Max, et al.
Publicado: (2021)
por: Hopkins, Max, et al.
Publicado: (2021)
Attention is All You Need Until You Need Retention
por: Yaslioglu, M. Murat
Publicado: (2025)
por: Yaslioglu, M. Murat
Publicado: (2025)
Standard Gaussian Process is All You Need for High-Dimensional Bayesian Optimization
por: Xu, Zhitong, et al.
Publicado: (2024)
por: Xu, Zhitong, et al.
Publicado: (2024)
Data is All You Need: Markov Chain Car-Following (MC-CF) Model
por: Chung, Sungyong, et al.
Publicado: (2026)
por: Chung, Sungyong, et al.
Publicado: (2026)
Context is All You Need
por: Delanois, Jean Erik, et al.
Publicado: (2026)
por: Delanois, Jean Erik, et al.
Publicado: (2026)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
por: Tyukin, Georgy, et al.
Publicado: (2024)
por: Tyukin, Georgy, et al.
Publicado: (2024)
Provable Privacy Attacks on Trained Shallow Neural Networks
por: Smorodinsky, Guy, et al.
Publicado: (2024)
por: Smorodinsky, Guy, et al.
Publicado: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
por: Chmiel, Brian, et al.
Publicado: (2025)
por: Chmiel, Brian, et al.
Publicado: (2025)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
por: Haroush, Matan, et al.
Publicado: (2025)
por: Haroush, Matan, et al.
Publicado: (2025)
Recursive Models for Long-Horizon Reasoning
por: Yang, Chenxiao, et al.
Publicado: (2026)
por: Yang, Chenxiao, et al.
Publicado: (2026)
Some Attention is All You Need for Retrieval
por: Michalak, Felix, et al.
Publicado: (2025)
por: Michalak, Felix, et al.
Publicado: (2025)
Half Search Space is All You Need
por: Rumiantsev, Pavel, et al.
Publicado: (2025)
por: Rumiantsev, Pavel, et al.
Publicado: (2025)
Top-$nσ$: Not All Logits Are You Need
por: Tang, Chenxia, et al.
Publicado: (2024)
por: Tang, Chenxia, et al.
Publicado: (2024)
Isoperimetry is All We Need: Langevin Posterior Sampling for RL with Sublinear Regret
por: Jorge, Emilio, et al.
Publicado: (2024)
por: Jorge, Emilio, et al.
Publicado: (2024)
Workspace Optimization: How to Train Your Agent
por: Sarafian, Elad, et al.
Publicado: (2026)
por: Sarafian, Elad, et al.
Publicado: (2026)
Exploitation Is All You Need... for Exploration
por: Rentschler, Micah, et al.
Publicado: (2025)
por: Rentschler, Micah, et al.
Publicado: (2025)
Multistep Inverse Is Not All You Need
por: Levine, Alexander, et al.
Publicado: (2024)
por: Levine, Alexander, et al.
Publicado: (2024)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
por: Joshi, Nirmit, et al.
Publicado: (2024)
por: Joshi, Nirmit, et al.
Publicado: (2024)
Ejemplares similares
-
Provable Tempered Overfitting of Minimal Nets and Typical Nets
por: Harel, Itamar, et al.
Publicado: (2024) -
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
por: Buzaglo, Gon, et al.
Publicado: (2024) -
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
por: Joshi, Nirmit, et al.
Publicado: (2023) -
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
por: Medvedev, Marko, et al.
Publicado: (2024) -
On the Hardness of Learning Regular Expressions
por: Attias, Idan, et al.
Publicado: (2025)