Gespeichert in:
| Hauptverfasser: | Nitanda, Atsushi, Kikuchi, Ryuhei, Maeda, Shugo, Wu, Denny |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2302.09376 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uniform convergence of the smooth calibration error and its relationship with functional gradient
von: Futami, Futoshi, et al.
Veröffentlicht: (2025)
von: Futami, Futoshi, et al.
Veröffentlicht: (2025)
Improved Particle Approximation Error for Mean Field Neural Networks
von: Nitanda, Atsushi
Veröffentlicht: (2024)
von: Nitanda, Atsushi
Veröffentlicht: (2024)
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
von: Maeda, Ibuki, et al.
Veröffentlicht: (2025)
von: Maeda, Ibuki, et al.
Veröffentlicht: (2025)
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
von: Takagi, Hirohane, et al.
Veröffentlicht: (2026)
von: Takagi, Hirohane, et al.
Veröffentlicht: (2026)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
Mirror Descent Policy Optimisation for Robust Constrained Markov Decision Processes
von: Bossens, David M., et al.
Veröffentlicht: (2025)
von: Bossens, David M., et al.
Veröffentlicht: (2025)
Emergence and scaling laws in SGD learning of shallow neural networks
von: Ren, Yunwei, et al.
Veröffentlicht: (2025)
von: Ren, Yunwei, et al.
Veröffentlicht: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
von: Chen, Zonghao, et al.
Veröffentlicht: (2025)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
von: Kawata, Ryotaro, et al.
Veröffentlicht: (2025)
Slowly Annealed Langevin Dynamics: Theory and Applications to Training-Free Guided Generation
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2026)
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2026)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
von: Lee, Jason D., et al.
Veröffentlicht: (2024)
von: Lee, Jason D., et al.
Veröffentlicht: (2024)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
von: Fu, Guoji, et al.
Veröffentlicht: (2026)
von: Fu, Guoji, et al.
Veröffentlicht: (2026)
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
von: Arous, Gérard Ben, et al.
Veröffentlicht: (2025)
von: Arous, Gérard Ben, et al.
Veröffentlicht: (2025)
Why pre-training is beneficial for downstream classification tasks?
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
von: Jiang, Xin, et al.
Veröffentlicht: (2024)
Koopman-based generalization bound: New aspect for full-rank weights
von: Hashimoto, Yuka, et al.
Veröffentlicht: (2023)
von: Hashimoto, Yuka, et al.
Veröffentlicht: (2023)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
von: Kovačević, Filip, et al.
Veröffentlicht: (2026)
von: Kovačević, Filip, et al.
Veröffentlicht: (2026)
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2025)
von: Nitanda, Atsushi, et al.
Veröffentlicht: (2025)
SGD method for entropy error function with smoothing l0 regularization for neural networks
von: Nguyen, Trong-Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Tuan, et al.
Veröffentlicht: (2024)
Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning
von: Bu, Dake, et al.
Veröffentlicht: (2024)
von: Bu, Dake, et al.
Veröffentlicht: (2024)
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
von: Bu, Dake, et al.
Veröffentlicht: (2026)
von: Bu, Dake, et al.
Veröffentlicht: (2026)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
From Coupled Oscillators to Graph Neural Networks: Reducing Over-smoothing via a Kuramoto Model-based Approach
von: Nguyen, Tuan, et al.
Veröffentlicht: (2023)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2023)
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
von: Chen, Feng, et al.
Veröffentlicht: (2023)
von: Chen, Feng, et al.
Veröffentlicht: (2023)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
von: Tao, Hongyi, et al.
Veröffentlicht: (2026)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
Post-Training as Reweighting: A Stochastic View of Reasoning Trajectories in Language Models
von: Bu, Dake, et al.
Veröffentlicht: (2025)
von: Bu, Dake, et al.
Veröffentlicht: (2025)
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2023)
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2023)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
von: Marek, Martin, et al.
Veröffentlicht: (2025)
von: Marek, Martin, et al.
Veröffentlicht: (2025)
Accelerated zero-order SGD under high-order smoothness and overparameterized regime
von: Bychkov, Georgii, et al.
Veröffentlicht: (2024)
von: Bychkov, Georgii, et al.
Veröffentlicht: (2024)
Demystifying Why Local Aggregation Helps: Convergence Analysis of Hierarchical SGD
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
von: Wang, Jiayi, et al.
Veröffentlicht: (2020)
Propagation of Chaos in One-hidden-layer Neural Networks beyond Logarithmic Time
von: Glasgow, Margalit, et al.
Veröffentlicht: (2025)
von: Glasgow, Margalit, et al.
Veröffentlicht: (2025)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
von: Jin, Richeng, et al.
Veröffentlicht: (2020)
Ordered Momentum for Asynchronous SGD
von: Shi, Chang-Wei, et al.
Veröffentlicht: (2024)
von: Shi, Chang-Wei, et al.
Veröffentlicht: (2024)
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems
von: Su, Junwei, et al.
Veröffentlicht: (2024)
von: Su, Junwei, et al.
Veröffentlicht: (2024)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yiheng, et al.
Veröffentlicht: (2026)
Generalization and Optimization of SGD with Lookahead
von: Li, Kangcheng, et al.
Veröffentlicht: (2025)
von: Li, Kangcheng, et al.
Veröffentlicht: (2025)
PCDP-SGD: Improving the Convergence of Differentially Private SGD via Projection in Advance
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
von: Sha, Haichao, et al.
Veröffentlicht: (2023)
Smoothed SGD for quantiles: Bahadur representation and Gaussian approximation
von: Chen, Likai, et al.
Veröffentlicht: (2025)
von: Chen, Likai, et al.
Veröffentlicht: (2025)
Topology-aware Generalization of Decentralized SGD
von: Zhu, Tongtian, et al.
Veröffentlicht: (2022)
von: Zhu, Tongtian, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Uniform convergence of the smooth calibration error and its relationship with functional gradient
von: Futami, Futoshi, et al.
Veröffentlicht: (2025) -
Improved Particle Approximation Error for Mean Field Neural Networks
von: Nitanda, Atsushi
Veröffentlicht: (2024) -
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
von: Maeda, Ibuki, et al.
Veröffentlicht: (2025) -
Alternating Diffusion for Proximal Sampling with Zeroth Order Queries
von: Takagi, Hirohane, et al.
Veröffentlicht: (2026) -
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)