Improved state mixing in higher-order and block diagonal linear recurrent networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dubinin, Igor, Orvieto, Antonio, Effenberger, Felix |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fading memory as inductive bias in residual recurrent networks
von: Dubinin, Igor, et al.
Veröffentlicht: (2023)
von: Dubinin, Igor, et al.
Veröffentlicht: (2023)
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024)
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025)
von: Laing, Sam, et al.
Veröffentlicht: (2025)
Precise asymptotics of reweighted least-squares algorithms for linear diagonal networks
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2024)
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2024)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2024)
In Search of Adam's Secret Sauce
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2025)
Fixed-Point RNNs: Interpolating from Diagonal to Dense
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2025)
Bridging the Gap Between Climate Science and Machine Learning in Climate Model Emulation
von: Schmidt, Luca, et al.
Veröffentlicht: (2026)
von: Schmidt, Luca, et al.
Veröffentlicht: (2026)
Explaining Grokking in Transformers through the Lens of Inductive Bias
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
von: Singh, Jaisidh, et al.
Veröffentlicht: (2026)
Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
von: Belloni, Annalisa, et al.
Veröffentlicht: (2026)
An Uncertainty Principle for Linear Recurrent Neural Networks
von: François, Alexandre, et al.
Veröffentlicht: (2025)
von: François, Alexandre, et al.
Veröffentlicht: (2025)
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
Turbine location-aware multi-decadal wind power predictions for Germany using CMIP6
von: Effenberger, Nina, et al.
Veröffentlicht: (2024)
von: Effenberger, Nina, et al.
Veröffentlicht: (2024)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
von: Srećković, Teodora, et al.
Veröffentlicht: (2025)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
von: Orvieto, Antonio, et al.
Veröffentlicht: (2023)
Gradient-free training of recurrent neural networks
von: Bolager, Erik Lien, et al.
Veröffentlicht: (2024)
von: Bolager, Erik Lien, et al.
Veröffentlicht: (2024)
Dance recalibration for dance coherency with recurrent convolution block
von: Eum, Seungho, et al.
Veröffentlicht: (2025)
von: Eum, Seungho, et al.
Veröffentlicht: (2025)
Towards Understanding Self-Pretraining for Sequence Classification
von: Coser, Omar, et al.
Veröffentlicht: (2026)
von: Coser, Omar, et al.
Veröffentlicht: (2026)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
von: Movahedi, Sajad, et al.
Veröffentlicht: (2024)
Super Consistency of Neural Network Landscapes and Learning Rate Transfer
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
von: Noci, Lorenzo, et al.
Veröffentlicht: (2024)
Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias
von: Zhang, Shuofeng, et al.
Veröffentlicht: (2025)
von: Zhang, Shuofeng, et al.
Veröffentlicht: (2025)
Exploring higher-order neural network node interactions with total correlation
von: Kerby, Thomas, et al.
Veröffentlicht: (2024)
von: Kerby, Thomas, et al.
Veröffentlicht: (2024)
NIMBA: Towards Robust and Principled Processing of Point Clouds With SSMs
von: Köprücü, Nursena, et al.
Veröffentlicht: (2024)
von: Köprücü, Nursena, et al.
Veröffentlicht: (2024)
On the low-shot transferability of [V]-Mamba
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
von: Misra, Diganta, et al.
Veröffentlicht: (2024)
Geometric sparsification in recurrent neural networks
von: Mackey, Wyatt, et al.
Veröffentlicht: (2024)
von: Mackey, Wyatt, et al.
Veröffentlicht: (2024)
How noise affects memory in linear recurrent networks
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
von: Guan, JingChuan, et al.
Veröffentlicht: (2024)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
Muown: Row-Norm Control for Muon Optimization
von: Lion, Kai, et al.
Veröffentlicht: (2026)
von: Lion, Kai, et al.
Veröffentlicht: (2026)
Quaternion recurrent neural network with real-time recurrent learning and maximum correntropy criterion
von: Bourigault, Pauline, et al.
Veröffentlicht: (2024)
von: Bourigault, Pauline, et al.
Veröffentlicht: (2024)
Recurrent Distance Filtering for Graph Representation Learning
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
von: Ding, Yuhui, et al.
Veröffentlicht: (2023)
Inferring stochastic low-rank recurrent neural networks from neural data
von: Pals, Matthijs, et al.
Veröffentlicht: (2024)
von: Pals, Matthijs, et al.
Veröffentlicht: (2024)
Downscaling land surface temperature data using edge detection and block-diagonal Gaussian process regression
von: Dandapanthula, Sanjit, et al.
Veröffentlicht: (2026)
von: Dandapanthula, Sanjit, et al.
Veröffentlicht: (2026)
Memory of recurrent networks: Do we compute it right?
von: Ballarin, Giovanni, et al.
Veröffentlicht: (2023)
von: Ballarin, Giovanni, et al.
Veröffentlicht: (2023)
An analog-electronic implementation of a harmonic oscillator recurrent neural network
von: Carvalho, Pedro, et al.
Veröffentlicht: (2025)
von: Carvalho, Pedro, et al.
Veröffentlicht: (2025)
GASP: Guided Asymmetric Self-Play For Coding LLMs
von: Jana, Swadesh, et al.
Veröffentlicht: (2026)
von: Jana, Swadesh, et al.
Veröffentlicht: (2026)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
Design Principles for Sequence Models via Coefficient Dynamics
von: Sieber, Jerome, et al.
Veröffentlicht: (2025)
von: Sieber, Jerome, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fading memory as inductive bias in residual recurrent networks
von: Dubinin, Igor, et al.
Veröffentlicht: (2023) -
Revisiting associative recall in modern recurrent models
von: Okpekpe, Destiny, et al.
Veröffentlicht: (2025) -
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2024) -
Adam Simplified: Bias Correction Debunked
von: Laing, Sam, et al.
Veröffentlicht: (2025) -
Precise asymptotics of reweighted least-squares algorithms for linear diagonal networks
von: Kaushik, Chiraag, et al.
Veröffentlicht: (2024)