Revisiting Glorot Initialization for Long-Range Linear Recurrences
Fuente:
arXiv
Saved in:
| Main Authors: | Bar, Noga, Seleznova, Mariia, Alexander, Yotam, Kutyniok, Gitta, Giryes, Raja |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
by: Seleznova, Mariia
Published: (2026)
by: Seleznova, Mariia
Published: (2026)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
by: Bar, Noga, et al.
Published: (2024)
by: Bar, Noga, et al.
Published: (2024)
Pruning at Initialization -- A Sketching Perspective
by: Bar, Noga, et al.
Published: (2023)
by: Bar, Noga, et al.
Published: (2023)
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
ZOQO: Zero-Order Quantized Optimization
by: Bar, Noga, et al.
Published: (2025)
by: Bar, Noga, et al.
Published: (2025)
GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection
by: Seleznova, Mariia, et al.
Published: (2025)
by: Seleznova, Mariia, et al.
Published: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
by: Boche, Holger, et al.
Published: (2024)
by: Boche, Holger, et al.
Published: (2024)
Overflow Prevention Enhances Long-Context Recurrent LLMs
by: Ben-Kish, Assaf, et al.
Published: (2025)
by: Ben-Kish, Assaf, et al.
Published: (2025)
Group Orthogonalization Regularization For Vision Models Adaptation and Robustness
by: Kurtz, Yoav, et al.
Published: (2023)
by: Kurtz, Yoav, et al.
Published: (2023)
Multiplicative Reweighting for Robust Neural Network Optimization
by: Bar, Noga, et al.
Published: (2021)
by: Bar, Noga, et al.
Published: (2021)
ParFam -- (Neural Guided) Symbolic Regression Based on Continuous Global Optimization
by: Scholl, Philipp, et al.
Published: (2023)
by: Scholl, Philipp, et al.
Published: (2023)
On the Relation Between Linear Diffusion and Power Iteration
by: Weitzner, Dana, et al.
Published: (2024)
by: Weitzner, Dana, et al.
Published: (2024)
Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
by: Mazza, Lorenzo, et al.
Published: (2026)
by: Mazza, Lorenzo, et al.
Published: (2026)
Sparse-Aware Neural Networks for Nonlinear Functionals: Mitigating the Exponential Dependence on Dimension
by: Li, Jianfei, et al.
Published: (2026)
by: Li, Jianfei, et al.
Published: (2026)
Graph Hierarchical Recurrence for Long-Range Generalization
by: Carotti, Stefano, et al.
Published: (2026)
by: Carotti, Stefano, et al.
Published: (2026)
Revisiting Long-term Time Series Forecasting: An Investigation on Linear Mapping
by: Li, Zhe, et al.
Published: (2023)
by: Li, Zhe, et al.
Published: (2023)
Entropy-Preserving Reinforcement Learning
by: Petrenko, Aleksei, et al.
Published: (2026)
by: Petrenko, Aleksei, et al.
Published: (2026)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
When is a System Discoverable from Data? Discovery Requires Chaos
by: Shumaylov, Zakhar, et al.
Published: (2025)
by: Shumaylov, Zakhar, et al.
Published: (2025)
Can Continuous-Time Diffusion Models Generate and Solve Globally Constrained Discrete Problems? A Study on Sudoku
by: Drozdova, Mariia
Published: (2026)
by: Drozdova, Mariia
Published: (2026)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
by: Ben-Kish, Assaf, et al.
Published: (2024)
by: Ben-Kish, Assaf, et al.
Published: (2024)
CARES: Context-Aware Resolution Selector for VLMs
by: Kimhi, Moshe, et al.
Published: (2025)
by: Kimhi, Moshe, et al.
Published: (2025)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
by: Norman, Yotam, et al.
Published: (2025)
by: Norman, Yotam, et al.
Published: (2025)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Linear Contextual Bandits with Hybrid Payoff: Revisited
by: Das, Nirjhar, et al.
Published: (2024)
by: Das, Nirjhar, et al.
Published: (2024)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
Probabilistic neural operators for functional uncertainty quantification
by: Bülte, Christopher, et al.
Published: (2025)
by: Bülte, Christopher, et al.
Published: (2025)
Random Spiking Neural Networks are Stable and Spectrally Simple
by: Araya, Ernesto, et al.
Published: (2025)
by: Araya, Ernesto, et al.
Published: (2025)
Generalization Bounds for Message Passing Networks on Mixture of Graphons
by: Maskey, Sohir, et al.
Published: (2024)
by: Maskey, Sohir, et al.
Published: (2024)
Long-Range Graph Wavelet Networks
by: Guerranti, Filippo, et al.
Published: (2025)
by: Guerranti, Filippo, et al.
Published: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
by: Zhao, Yike, et al.
Published: (2026)
by: Zhao, Yike, et al.
Published: (2026)
Optimal Decay Spectra for Linear Recurrences
by: Cao, Yang
Published: (2026)
by: Cao, Yang
Published: (2026)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Interpreting Affine Recurrence Learning in GPT-style Transformers
by: Bhargav, Samarth, et al.
Published: (2024)
by: Bhargav, Samarth, et al.
Published: (2024)
Block-Biased Mamba for Long-Range Sequence Processing
by: Yu, Annan, et al.
Published: (2025)
by: Yu, Annan, et al.
Published: (2025)
Provable Long-Range Benefits of Next-Token Prediction
by: Cao, Xinyuan, et al.
Published: (2025)
by: Cao, Xinyuan, et al.
Published: (2025)
On Measuring Long-Range Interactions in Graph Neural Networks
by: Bamberger, Jacob, et al.
Published: (2025)
by: Bamberger, Jacob, et al.
Published: (2025)
Similar Items
-
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
by: Seleznova, Mariia
Published: (2026) -
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
by: Bar, Noga, et al.
Published: (2024) -
Pruning at Initialization -- A Sketching Perspective
by: Bar, Noga, et al.
Published: (2023) -
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
by: Razin, Noam, et al.
Published: (2024) -
ZOQO: Zero-Order Quantized Optimization
by: Bar, Noga, et al.
Published: (2025)