Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Sato, Naoki, Iiduka, Hideaki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Explicit and Implicit Graduated Optimization in Deep Neural Networks
di: Sato, Naoki, et al.
Pubblicazione: (2024)
di: Sato, Naoki, et al.
Pubblicazione: (2024)
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
di: Sato, Naoki, et al.
Pubblicazione: (2023)
di: Sato, Naoki, et al.
Pubblicazione: (2023)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
di: Sato, Naoki, et al.
Pubblicazione: (2024)
di: Sato, Naoki, et al.
Pubblicazione: (2024)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
di: Sato, Naoki, et al.
Pubblicazione: (2024)
di: Sato, Naoki, et al.
Pubblicazione: (2024)
Convergence Bound and Critical Batch Size of Muon Optimizer
di: Sato, Naoki, et al.
Pubblicazione: (2025)
di: Sato, Naoki, et al.
Pubblicazione: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
di: Kondo, Yuichi, et al.
Pubblicazione: (2025)
di: Kondo, Yuichi, et al.
Pubblicazione: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
di: Umeda, Hikaru, et al.
Pubblicazione: (2024)
di: Umeda, Hikaru, et al.
Pubblicazione: (2024)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
di: Iiduka, Hideaki
Pubblicazione: (2026)
di: Iiduka, Hideaki
Pubblicazione: (2026)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
di: Imaizumi, Kento, et al.
Pubblicazione: (2024)
di: Imaizumi, Kento, et al.
Pubblicazione: (2024)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
di: Kamo, Keisuke, et al.
Pubblicazione: (2025)
di: Kamo, Keisuke, et al.
Pubblicazione: (2025)
Convergence Analysis of SGD under Expected Smoothness
di: Kawamoto, Yuta, et al.
Pubblicazione: (2025)
di: Kawamoto, Yuta, et al.
Pubblicazione: (2025)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
di: Nagashima, Shuntaro, et al.
Pubblicazione: (2026)
di: Nagashima, Shuntaro, et al.
Pubblicazione: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
di: Oowada, Kanata, et al.
Pubblicazione: (2025)
di: Oowada, Kanata, et al.
Pubblicazione: (2025)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
di: Tsukada, Yuki, et al.
Pubblicazione: (2023)
di: Tsukada, Yuki, et al.
Pubblicazione: (2023)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
di: Harada, Hinata, et al.
Pubblicazione: (2024)
di: Harada, Hinata, et al.
Pubblicazione: (2024)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
di: Imaizumi, Kento, et al.
Pubblicazione: (2025)
di: Imaizumi, Kento, et al.
Pubblicazione: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
di: Umeda, Hikaru, et al.
Pubblicazione: (2025)
A Nonparametric Statistics Approach to Feature Selection in Deep Neural Networks with Theoretical Guarantees
di: Du, Junye, et al.
Pubblicazione: (2025)
di: Du, Junye, et al.
Pubblicazione: (2025)
Bi-Lipschitz Autoencoder With Injectivity Guarantee
di: Zhan, Qipeng, et al.
Pubblicazione: (2026)
di: Zhan, Qipeng, et al.
Pubblicazione: (2026)
Geometry-Aware Approaches for Balancing Performance and Theoretical Guarantees in Linear Bandits
di: Luo, Yuwei, et al.
Pubblicazione: (2023)
di: Luo, Yuwei, et al.
Pubblicazione: (2023)
Theoretical Guarantees for Low-Rank Compression of Deep Neural Networks
di: Zhang, Shihao, et al.
Pubblicazione: (2025)
di: Zhang, Shihao, et al.
Pubblicazione: (2025)
Reversible Deep Equilibrium Models
di: McCallum, Sam, et al.
Pubblicazione: (2025)
di: McCallum, Sam, et al.
Pubblicazione: (2025)
Theoretically Guaranteed Distribution Adaptable Learning
di: Xu, Chao, et al.
Pubblicazione: (2024)
di: Xu, Chao, et al.
Pubblicazione: (2024)
Theoretical Convergence Guarantees for Variational Autoencoders
di: Surendran, Sobihan, et al.
Pubblicazione: (2024)
di: Surendran, Sobihan, et al.
Pubblicazione: (2024)
Mini-Batch Stochastic Halpern Algorithm for Nonexpansive Fixed point Problems
di: Iiduka, Hideaki
Pubblicazione: (2026)
di: Iiduka, Hideaki
Pubblicazione: (2026)
Mini-Batch Stochastic Krasnosel'ski\uı-Mann Algorithm for Nonexpansive Fixed Point Problems
di: Iiduka, Hideaki
Pubblicazione: (2026)
di: Iiduka, Hideaki
Pubblicazione: (2026)
Clustering with Tangles: Algorithmic Framework and Theoretical Guarantees
di: Klepper, Solveig, et al.
Pubblicazione: (2020)
di: Klepper, Solveig, et al.
Pubblicazione: (2020)
Online Time Series Forecasting with Theoretical Guarantees
di: Li, Zijian, et al.
Pubblicazione: (2025)
di: Li, Zijian, et al.
Pubblicazione: (2025)
Consistency Deep Equilibrium Models
di: Lin, Junchao, et al.
Pubblicazione: (2026)
di: Lin, Junchao, et al.
Pubblicazione: (2026)
Quantum Deep Equilibrium Models
di: Schleich, Philipp, et al.
Pubblicazione: (2024)
di: Schleich, Philipp, et al.
Pubblicazione: (2024)
Information Theoretic Guarantees For Policy Alignment In Large Language Models
di: Mroueh, Youssef
Pubblicazione: (2024)
di: Mroueh, Youssef
Pubblicazione: (2024)
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
di: Jin, Richeng, et al.
Pubblicazione: (2020)
di: Jin, Richeng, et al.
Pubblicazione: (2020)
Clustering-based Meta Bayesian Optimization with Theoretical Guarantee
di: Nguyen, Khoa, et al.
Pubblicazione: (2025)
di: Nguyen, Khoa, et al.
Pubblicazione: (2025)
Transformation-Invariant Learning and Theoretical Guarantees for OOD Generalization
di: Montasser, Omar, et al.
Pubblicazione: (2024)
di: Montasser, Omar, et al.
Pubblicazione: (2024)
Theoretical Guarantees for Causal Discovery on Large Random Graphs
di: Chevalley, Mathieu, et al.
Pubblicazione: (2025)
di: Chevalley, Mathieu, et al.
Pubblicazione: (2025)
Adaptive Initial Residual Connections for GNNs with Theoretical Guarantees
di: Shirzadi, Mohammad, et al.
Pubblicazione: (2025)
di: Shirzadi, Mohammad, et al.
Pubblicazione: (2025)
Operator Learning of Lipschitz Operators: An Information-Theoretic Perspective
di: Lanthaler, Samuel
Pubblicazione: (2024)
di: Lanthaler, Samuel
Pubblicazione: (2024)
Convergence Guarantees for the DeepWalk Embedding on Block Models
di: Harker, Christopher, et al.
Pubblicazione: (2024)
di: Harker, Christopher, et al.
Pubblicazione: (2024)
A Generalized Meta Federated Learning Framework with Theoretical Convergence Guarantees
di: Jamali, Mohammad Vahid, et al.
Pubblicazione: (2025)
di: Jamali, Mohammad Vahid, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Explicit and Implicit Graduated Optimization in Deep Neural Networks
di: Sato, Naoki, et al.
Pubblicazione: (2024) -
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
di: Sato, Naoki, et al.
Pubblicazione: (2023) -
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
di: Sato, Naoki, et al.
Pubblicazione: (2024) -
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
di: Sato, Naoki, et al.
Pubblicazione: (2024) -
Convergence Bound and Critical Batch Size of Muon Optimizer
di: Sato, Naoki, et al.
Pubblicazione: (2025)