Recurrent neural networks: vanishing and exploding gradients are not the end of the story
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zucchet, Nicolas, Orvieto, Antonio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
von: Grontas, Panagiotis D., et al.
Veröffentlicht: (2025)
von: Grontas, Panagiotis D., et al.
Veröffentlicht: (2025)
Precise gradient descent training dynamics for finite-width multi-layer neural networks
von: Han, Qiyang, et al.
Veröffentlicht: (2025)
von: Han, Qiyang, et al.
Veröffentlicht: (2025)
A multiobjective continuation method to compute the regularization path of deep neural networks
von: Amakor, Augustina C., et al.
Veröffentlicht: (2023)
von: Amakor, Augustina C., et al.
Veröffentlicht: (2023)
Explicit neural network classifiers for non-separable data
von: Ewald, Patrícia Muñoz
Veröffentlicht: (2025)
von: Ewald, Patrícia Muñoz
Veröffentlicht: (2025)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
When majority rules, minority loses: bias amplification of gradient descent
von: Bachoc, François, et al.
Veröffentlicht: (2025)
von: Bachoc, François, et al.
Veröffentlicht: (2025)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
von: Shen, Jucheng, et al.
Veröffentlicht: (2026)
von: Shen, Jucheng, et al.
Veröffentlicht: (2026)
A second-order-like optimizer with adaptive gradient scaling for deep learning
von: Bolte, Jérôme, et al.
Veröffentlicht: (2024)
von: Bolte, Jérôme, et al.
Veröffentlicht: (2024)
Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity
von: Yang, Yan, et al.
Veröffentlicht: (2024)
von: Yang, Yan, et al.
Veröffentlicht: (2024)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2022)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2022)
Convergence of gradient flow for learning convolutional neural networks
von: Diederen, Jona-Maria, et al.
Veröffentlicht: (2026)
von: Diederen, Jona-Maria, et al.
Veröffentlicht: (2026)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
Frequency-aware Surrogate Modeling With SMT Kernels For Advanced Data Forecasting
von: Gonel, Nicolas, et al.
Veröffentlicht: (2025)
von: Gonel, Nicolas, et al.
Veröffentlicht: (2025)
The duality structure gradient descent algorithm: analysis and applications to neural networks
von: Flynn, Thomas
Veröffentlicht: (2017)
von: Flynn, Thomas
Veröffentlicht: (2017)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
von: Choudhury, Sayantan, et al.
Veröffentlicht: (2024)
von: Choudhury, Sayantan, et al.
Veröffentlicht: (2024)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
von: Lugosi, Gabor, et al.
Veröffentlicht: (2024)
von: Lugosi, Gabor, et al.
Veröffentlicht: (2024)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
The Algorithm Configuration Problem
von: Iommazzo, Gabriele, et al.
Veröffentlicht: (2024)
von: Iommazzo, Gabriele, et al.
Veröffentlicht: (2024)
Understanding the differences in Foundation Models: Attention, State Space Models, and Recurrent Neural Networks
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
von: Sieber, Jerome, et al.
Veröffentlicht: (2024)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
von: An, Jing, et al.
Veröffentlicht: (2023)
von: An, Jing, et al.
Veröffentlicht: (2023)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
von: Pacheco, Bruno Machado, et al.
Veröffentlicht: (2023)
von: Pacheco, Bruno Machado, et al.
Veröffentlicht: (2023)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
MIQCQP reformulation of the ReLU neural networks Lipschitz constant estimation problem
von: Sbihi, Mohammed, et al.
Veröffentlicht: (2024)
von: Sbihi, Mohammed, et al.
Veröffentlicht: (2024)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
Reward-Directed Score-Based Diffusion Models via q-Learning
von: Gao, Xuefeng, et al.
Veröffentlicht: (2024)
von: Gao, Xuefeng, et al.
Veröffentlicht: (2024)
Painless Federated Learning: An Interplay of Line-Search and Extrapolation
von: Geetika, et al.
Veröffentlicht: (2024)
von: Geetika, et al.
Veröffentlicht: (2024)
Understanding Optimization in Deep Learning with Central Flows
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
von: Cohen, Jeremy M., et al.
Veröffentlicht: (2024)
Adaptive Primal-Dual Method for Safe Reinforcement Learning
von: Chen, Weiqin, et al.
Veröffentlicht: (2024)
von: Chen, Weiqin, et al.
Veröffentlicht: (2024)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
von: Wasserkrug, Segev, et al.
Veröffentlicht: (2024)
von: Wasserkrug, Segev, et al.
Veröffentlicht: (2024)
Training Safe Neural Networks with Global SDP Bounds
von: Soletskyi, Roman, et al.
Veröffentlicht: (2024)
von: Soletskyi, Roman, et al.
Veröffentlicht: (2024)
Boosting Gradient Ascent for Continuous DR-submodular Maximization
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
von: Zhang, Qixin, et al.
Veröffentlicht: (2024)
Learning Backdoors for Mixed Integer Linear Programs with Contrastive Learning
von: Cai, Junyang, et al.
Veröffentlicht: (2024)
von: Cai, Junyang, et al.
Veröffentlicht: (2024)
Generating Likely Counterfactuals Using Sum-Product Networks
von: Nemecek, Jiri, et al.
Veröffentlicht: (2024)
von: Nemecek, Jiri, et al.
Veröffentlicht: (2024)
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack
von: Gai, Kuo, et al.
Veröffentlicht: (2024)
von: Gai, Kuo, et al.
Veröffentlicht: (2024)
An Efficient Learning-based Solver Comparable to Metaheuristics for the Capacitated Arc Routing Problem
von: Guo, Runze, et al.
Veröffentlicht: (2024)
von: Guo, Runze, et al.
Veröffentlicht: (2024)
A Benchmark for Maximum Cut: Towards Standardization of the Evaluation of Learned Heuristics for Combinatorial Optimization
von: Nath, Ankur, et al.
Veröffentlicht: (2024)
von: Nath, Ankur, et al.
Veröffentlicht: (2024)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
von: Hong, Seong-Hyun, et al.
Veröffentlicht: (2024)
von: Hong, Seong-Hyun, et al.
Veröffentlicht: (2024)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient and provably convergent end-to-end training of deep neural networks with linear constraints
von: Yang, Zonglin, et al.
Veröffentlicht: (2026) -
Pinet: Optimizing hard-constrained neural networks with orthogonal projection layers
von: Grontas, Panagiotis D., et al.
Veröffentlicht: (2025) -
Precise gradient descent training dynamics for finite-width multi-layer neural networks
von: Han, Qiyang, et al.
Veröffentlicht: (2025) -
A multiobjective continuation method to compute the regularization path of deep neural networks
von: Amakor, Augustina C., et al.
Veröffentlicht: (2023) -
Explicit neural network classifiers for non-separable data
von: Ewald, Patrícia Muñoz
Veröffentlicht: (2025)