On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Zhong, Han, Jiequn, E, Weinan, Li, Qianxiao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
por: Mitchell, Rupert, et al.
Publicado: (2025)
por: Mitchell, Rupert, et al.
Publicado: (2025)
Chaos-Free Networks are Stable Recurrent Neural Networks
por: De Carli, Stefano, et al.
Publicado: (2026)
por: De Carli, Stefano, et al.
Publicado: (2026)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
por: Cheridito, Patrick, et al.
Publicado: (2022)
por: Cheridito, Patrick, et al.
Publicado: (2022)
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
por: Alexander, Michael
Publicado: (2026)
por: Alexander, Michael
Publicado: (2026)
Task-Synchronized Recurrent Neural Networks
por: Lukoševičius, Mantas, et al.
Publicado: (2022)
por: Lukoševičius, Mantas, et al.
Publicado: (2022)
Stability properties of Minimal Gated Unit neural networks
por: De Carli, Stefano, et al.
Publicado: (2026)
por: De Carli, Stefano, et al.
Publicado: (2026)
A Domain Decomposition-Based CNN-DNN Architecture for Model Parallel Training Applied to Image Recognition Problems
por: Klawonn, Axel, et al.
Publicado: (2023)
por: Klawonn, Axel, et al.
Publicado: (2023)
NeurOptimisation: The Spiking Way to Evolve
por: Cruz-Duarte, Jorge Mario, et al.
Publicado: (2025)
por: Cruz-Duarte, Jorge Mario, et al.
Publicado: (2025)
DDU-Net: A Domain Decomposition-Based CNN for High-Resolution Image Segmentation on Multiple GPUs
por: Verburg, Corné, et al.
Publicado: (2024)
por: Verburg, Corné, et al.
Publicado: (2024)
Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis
por: Hennadii, Kutomanov
Publicado: (2026)
por: Hennadii, Kutomanov
Publicado: (2026)
Velocity-Inferred Hamiltonian Neural Networks: Learning Energy-Conserving Dynamics from Position-Only Data
por: Xu, Ruichen, et al.
Publicado: (2025)
por: Xu, Ruichen, et al.
Publicado: (2025)
Predicting Music Track Popularity by Convolutional Neural Networks on Spotify Features and Spectrogram of Audio Waveform
por: Falah, Navid, et al.
Publicado: (2025)
por: Falah, Navid, et al.
Publicado: (2025)
Approximating k-Center via Farthest-First on $δ$-Covers
por: Wilson, Jason R.
Publicado: (2026)
por: Wilson, Jason R.
Publicado: (2026)
Neural Encoding for Image Recall: Human-Like Memory
por: Foussereau, Virgile, et al.
Publicado: (2024)
por: Foussereau, Virgile, et al.
Publicado: (2024)
i-DEQ: A stable inertial deep equilibrium model for image restoration
por: Clerc, Antonin, et al.
Publicado: (2026)
por: Clerc, Antonin, et al.
Publicado: (2026)
Socio-cognitive agent-oriented evolutionary algorithm with trust-based optimization
por: Urbańczyk, Aleksandra, et al.
Publicado: (2025)
por: Urbańczyk, Aleksandra, et al.
Publicado: (2025)
Optimization approaches to Wolbachia-based biocontrol
por: Orozco-Gonzales, Jose L., et al.
Publicado: (2024)
por: Orozco-Gonzales, Jose L., et al.
Publicado: (2024)
Constrained Neural Networks for Interpretable Heuristic Creation to Optimise Computer Algebra Systems
por: Florescu, Dorian, et al.
Publicado: (2024)
por: Florescu, Dorian, et al.
Publicado: (2024)
Repetition Makes Perfect: Recurrent Graph Neural Networks Match Message-Passing Limit
por: Rosenbluth, Eran, et al.
Publicado: (2025)
por: Rosenbluth, Eran, et al.
Publicado: (2025)
Conservation Law Breaking at the Edge of Stability: A Spectral Theory of Non-Convex Neural Network Optimization
por: Medeiros, Daniel Nobrega
Publicado: (2026)
por: Medeiros, Daniel Nobrega
Publicado: (2026)
Prediction-space knowledge markets for communication-efficient federated learning on multimedia tasks
por: Du, Wenzhang
Publicado: (2025)
por: Du, Wenzhang
Publicado: (2025)
Stochastic Mirror Descent for Convex Optimization with Consensus Constraints
por: Borovykh, Anastasia, et al.
Publicado: (2022)
por: Borovykh, Anastasia, et al.
Publicado: (2022)
On the Origin of Algorithmic Progress in AI
por: Gundlach, Hans, et al.
Publicado: (2025)
por: Gundlach, Hans, et al.
Publicado: (2025)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
por: Brännvall, Rickard, et al.
Publicado: (2023)
por: Brännvall, Rickard, et al.
Publicado: (2023)
Convergence of Momentum-Based Optimization Algorithms with Time-Varying Parameters
por: Vidyasagar, Mathukumalli
Publicado: (2025)
por: Vidyasagar, Mathukumalli
Publicado: (2025)
TACIT: Transformation-Aware Capturing of Implicit Thought
por: Nobrega, Daniel
Publicado: (2026)
por: Nobrega, Daniel
Publicado: (2026)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
por: Piao, Yangheran, et al.
Publicado: (2025)
por: Piao, Yangheran, et al.
Publicado: (2025)
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
por: Zhuang, Yiyu, et al.
Publicado: (2024)
por: Zhuang, Yiyu, et al.
Publicado: (2024)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
por: Nayak, Prabhudarshi, et al.
Publicado: (2026)
por: Nayak, Prabhudarshi, et al.
Publicado: (2026)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
por: Cui, Hejie, et al.
Publicado: (2024)
por: Cui, Hejie, et al.
Publicado: (2024)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
por: Yamchote, Phaphontee, et al.
Publicado: (2025)
por: Yamchote, Phaphontee, et al.
Publicado: (2025)
Model-Free Local Recalibration of Neural Networks
por: Torres, R., et al.
Publicado: (2024)
por: Torres, R., et al.
Publicado: (2024)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
por: Zhang, Minxin, et al.
Publicado: (2026)
por: Zhang, Minxin, et al.
Publicado: (2026)
A Constraint-Preserving Neural Network Approach for Solving Mean-Field Games Equilibrium
por: Liu, Jinwei, et al.
Publicado: (2025)
por: Liu, Jinwei, et al.
Publicado: (2025)
PyEPO: A PyTorch-based End-to-End Predict-then-Optimize Library for Linear and Integer Programming
por: Tang, Bo, et al.
Publicado: (2022)
por: Tang, Bo, et al.
Publicado: (2022)
Multiple data-driven missing imputation
por: Kavun, Sergii
Publicado: (2025)
por: Kavun, Sergii
Publicado: (2025)
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
por: Danieli, Federico, et al.
Publicado: (2025)
por: Danieli, Federico, et al.
Publicado: (2025)
GLL: A Differentiable Graph Learning Layer for Neural Networks
por: Brown, Jason, et al.
Publicado: (2024)
por: Brown, Jason, et al.
Publicado: (2024)
Revisiting Non-separable Binary Classification and its Applications in Anomaly Detection
por: Lau, Matthew, et al.
Publicado: (2023)
por: Lau, Matthew, et al.
Publicado: (2023)
Neural Networks for Tamed Milstein Approximation of SDEs with Additive Symmetric Jump Noise Driven by a Poisson Random Measure
por: Ramirez-Gonzalez, Jose-Hermenegildo, et al.
Publicado: (2025)
por: Ramirez-Gonzalez, Jose-Hermenegildo, et al.
Publicado: (2025)
Ejemplares similares
-
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
por: Mitchell, Rupert, et al.
Publicado: (2025) -
Chaos-Free Networks are Stable Recurrent Neural Networks
por: De Carli, Stefano, et al.
Publicado: (2026) -
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
por: Cheridito, Patrick, et al.
Publicado: (2022) -
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
por: Alexander, Michael
Publicado: (2026) -
Task-Synchronized Recurrent Neural Networks
por: Lukoševičius, Mantas, et al.
Publicado: (2022)