A Quantitative Characterization of Forgetting in Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Balasubramanian, Krishnakumar, Kasiviswanathan, Shiva Prasad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
by: Balasubramanian, Krishnakumar
Published: (2026)
by: Balasubramanian, Krishnakumar
Published: (2026)
Dense associative memory for Gaussian distributions
by: Tankala, Chandan, et al.
Published: (2025)
by: Tankala, Chandan, et al.
Published: (2025)
Total Variation Rates for Riemannian Flow Matching
by: Guan, Yunrui, et al.
Published: (2026)
by: Guan, Yunrui, et al.
Published: (2026)
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression
by: Liang, Haodong, et al.
Published: (2025)
by: Liang, Haodong, et al.
Published: (2025)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
by: Park, Junhyung, et al.
Published: (2024)
by: Park, Junhyung, et al.
Published: (2024)
Sequential Kernelized Independence Testing
by: Podkopaev, Aleksandr, et al.
Published: (2022)
by: Podkopaev, Aleksandr, et al.
Published: (2022)
Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
by: Jia, Sheng, et al.
Published: (2025)
by: Jia, Sheng, et al.
Published: (2025)
Optimal Transportation and Alignment Between Gaussian Measures
by: Dandapanthula, Sanjit, et al.
Published: (2025)
by: Dandapanthula, Sanjit, et al.
Published: (2025)
Finite-Dimensional Gaussian Approximation for Deep Neural Networks: Universality in Random Weights
by: Balasubramanian, Krishnakumar, et al.
Published: (2025)
by: Balasubramanian, Krishnakumar, et al.
Published: (2025)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Large-Step Training Dynamics of a Two-Factor Linear Transformer Model
by: Balasubramanian, Krishnakumar
Published: (2026)
by: Balasubramanian, Krishnakumar
Published: (2026)
Riemannian Proximal Sampler for High-accuracy Sampling on Manifolds
by: Guan, Yunrui, et al.
Published: (2025)
by: Guan, Yunrui, et al.
Published: (2025)
Meta-Learning with Generalized Ridge Regression: High-dimensional Asymptotics, Optimality and Hyper-covariance Estimation
by: Jin, Yanhao, et al.
Published: (2024)
by: Jin, Yanhao, et al.
Published: (2024)
Statistical Inference for Linear Functionals of Online Least-squares SGD when $t \gtrsim d^{1+δ}$
by: Agrawalla, Bhavya, et al.
Published: (2025)
by: Agrawalla, Bhavya, et al.
Published: (2025)
Nonsmooth Nonparametric Regression via Fractional Laplacian Eigenmaps
by: Shi, Zhaoyang, et al.
Published: (2024)
by: Shi, Zhaoyang, et al.
Published: (2024)
The Coverage Principle: How Pre-Training Enables Post-Training
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Anytime-Valid Inference for Double/Debiased Machine Learning of Causal Parameters
by: Dalal, Abhinandan, et al.
Published: (2024)
by: Dalal, Abhinandan, et al.
Published: (2024)
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent
by: Banerjee, Sayan, et al.
Published: (2024)
by: Banerjee, Sayan, et al.
Published: (2024)
Finite-Particle Rates for Regularized Stein Variational Gradient Descent
by: He, Ye, et al.
Published: (2026)
by: He, Ye, et al.
Published: (2026)
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
by: He, Ye, et al.
Published: (2024)
by: He, Ye, et al.
Published: (2024)
Training Implicit Generative Models via an Invariant Statistical Loss
by: de Frutos, José Manuel, et al.
Published: (2024)
by: de Frutos, José Manuel, et al.
Published: (2024)
Multivariate Gaussian Approximation for Random Forest via Region-based Stabilization
by: Shi, Zhaoyang, et al.
Published: (2024)
by: Shi, Zhaoyang, et al.
Published: (2024)
Gaussian random field approximation via Stein's method with applications to wide random neural networks
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
by: Agrawalla, Bhavya, et al.
Published: (2023)
by: Agrawalla, Bhavya, et al.
Published: (2023)
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
by: Balasubramanian, Krishnakumar, et al.
Published: (2023)
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
by: Wang, Kevin, et al.
Published: (2026)
by: Wang, Kevin, et al.
Published: (2026)
Gaussian and Bootstrap Approximation for Matching-based Average Treatment Effect Estimators
by: Shi, Zhaoyang, et al.
Published: (2024)
by: Shi, Zhaoyang, et al.
Published: (2024)
What Causes Postoperative Aspiration?
by: Nagesh, Supriya, et al.
Published: (2025)
by: Nagesh, Supriya, et al.
Published: (2025)
Statistical inference with belief functions: A survey
by: Cuzzolin, Fabio
Published: (2026)
by: Cuzzolin, Fabio
Published: (2026)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
by: Kontorovich, Aryeh, et al.
Published: (2026)
by: Kontorovich, Aryeh, et al.
Published: (2026)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Deep Ensembles for Epistemic Uncertainty: A Frequentist Perspective
by: Jain, Anchit, et al.
Published: (2025)
by: Jain, Anchit, et al.
Published: (2025)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
by: Yan, Hedong
Published: (2025)
by: Yan, Hedong
Published: (2025)
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
by: Roy, Saptarshi, et al.
Published: (2026)
by: Roy, Saptarshi, et al.
Published: (2026)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
by: Hazard, Christopher J., et al.
Published: (2025)
by: Hazard, Christopher J., et al.
Published: (2025)
A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
Similar Items
-
Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
by: Balasubramanian, Krishnakumar
Published: (2026) -
Dense associative memory for Gaussian distributions
by: Tankala, Chandan, et al.
Published: (2025) -
Total Variation Rates for Riemannian Flow Matching
by: Guan, Yunrui, et al.
Published: (2026) -
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024) -
Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression
by: Liang, Haodong, et al.
Published: (2025)