Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yilan, Wang, Zhichao, Huang, Wei, Han, Andi, Suzuki, Taiji, Mazumdar, Arya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025)
by: Jiang, Jiarui, et al.
Published: (2025)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
by: Takakura, Shokichi, et al.
Published: (2024)
by: Takakura, Shokichi, et al.
Published: (2024)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization
by: Chen, Yihang, et al.
Published: (2024)
by: Chen, Yihang, et al.
Published: (2024)
Learning and Generalization with Mixture Data
by: Vardhan, Harsh, et al.
Published: (2025)
by: Vardhan, Harsh, et al.
Published: (2025)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
by: Zhang, Tongcheng, et al.
Published: (2026)
by: Zhang, Tongcheng, et al.
Published: (2026)
On the Theory of Continual Learning with Gradient Descent for Neural Networks
by: Taheri, Hossein, et al.
Published: (2025)
by: Taheri, Hossein, et al.
Published: (2025)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
by: Taheri, Hossein, et al.
Published: (2024)
by: Taheri, Hossein, et al.
Published: (2024)
On the Role of Label Noise in the Feature Learning Process
by: Han, Andi, et al.
Published: (2025)
by: Han, Andi, et al.
Published: (2025)
Deep Two-Way Matrix Reordering for Relational Data Analysis
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Zero-Flow Encoders
by: Wang, Yakun, et al.
Published: (2026)
by: Wang, Yakun, et al.
Published: (2026)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
by: Bu, Dake, et al.
Published: (2024)
by: Bu, Dake, et al.
Published: (2024)
Exact Recovery of Sparse Binary Vectors from Generalized Linear Measurements
by: Mazumdar, Arya, et al.
Published: (2025)
by: Mazumdar, Arya, et al.
Published: (2025)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
by: Awano, Ryoya, et al.
Published: (2026)
by: Awano, Ryoya, et al.
Published: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
by: Wakayama, Tomoya, et al.
Published: (2025)
by: Wakayama, Tomoya, et al.
Published: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
by: Oh, Junsoo, et al.
Published: (2025)
by: Oh, Junsoo, et al.
Published: (2025)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Effectiveness of Distributed Gradient Descent with Local Steps for Overparameterized Models
by: Zhu, Heng, et al.
Published: (2024)
by: Zhu, Heng, et al.
Published: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
by: Bu, Dake, et al.
Published: (2026)
by: Bu, Dake, et al.
Published: (2026)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
Collaborative Compressors in Distributed Mean Estimation with Limited Communication Budget
by: Vardhan, Harsh, et al.
Published: (2026)
by: Vardhan, Harsh, et al.
Published: (2026)
Test time training enhances in-context learning of nonlinear functions
by: Kuwataka, Kento, et al.
Published: (2025)
by: Kuwataka, Kento, et al.
Published: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
by: Nishikawa, Naoki, et al.
Published: (2024)
by: Nishikawa, Naoki, et al.
Published: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
by: Kim, Juno, et al.
Published: (2024)
by: Kim, Juno, et al.
Published: (2024)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
by: Watanabe, Chihiro, et al.
Published: (2021)
by: Watanabe, Chihiro, et al.
Published: (2021)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
by: Takakura, Shokichi, et al.
Published: (2023)
by: Takakura, Shokichi, et al.
Published: (2023)
Client Selection in Federated Learning with Data Heterogeneity and Network Latencies
by: Vardhan, Harsh, et al.
Published: (2025)
by: Vardhan, Harsh, et al.
Published: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
by: Chen, Zonghao, et al.
Published: (2025)
by: Chen, Zonghao, et al.
Published: (2025)
On the sample complexity of parameter estimation in logistic regression with normal design
by: Hsu, Daniel, et al.
Published: (2023)
by: Hsu, Daniel, et al.
Published: (2023)
Agnostic Learning of Mixed Linear Regressions with EM and AM Algorithms
by: Ghosh, Avishek, et al.
Published: (2024)
by: Ghosh, Avishek, et al.
Published: (2024)
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
by: Kawata, Ryotaro, et al.
Published: (2025)
by: Kawata, Ryotaro, et al.
Published: (2025)
Similar Items
-
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
by: Huang, Wei, et al.
Published: (2025) -
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024) -
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
by: Bu, Dake, et al.
Published: (2025) -
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
by: Jiang, Jiarui, et al.
Published: (2025) -
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024)