Revisiting Glorot Initialization for Long-Range Linear Recurrences
Fuente:
arXiv
Guardado en:
| Autores principales: | Bar, Noga, Seleznova, Mariia, Alexander, Yotam, Kutyniok, Gitta, Giryes, Raja |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
por: Seleznova, Mariia
Publicado: (2026)
por: Seleznova, Mariia
Publicado: (2026)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
por: Bar, Noga, et al.
Publicado: (2024)
por: Bar, Noga, et al.
Publicado: (2024)
Pruning at Initialization -- A Sketching Perspective
por: Bar, Noga, et al.
Publicado: (2023)
por: Bar, Noga, et al.
Publicado: (2023)
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
por: Razin, Noam, et al.
Publicado: (2024)
por: Razin, Noam, et al.
Publicado: (2024)
ZOQO: Zero-Order Quantized Optimization
por: Bar, Noga, et al.
Publicado: (2025)
por: Bar, Noga, et al.
Publicado: (2025)
GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection
por: Seleznova, Mariia, et al.
Publicado: (2025)
por: Seleznova, Mariia, et al.
Publicado: (2025)
Mathematical Algorithm Design for Deep Learning under Societal and Judicial Constraints: The Algorithmic Transparency Requirement
por: Boche, Holger, et al.
Publicado: (2024)
por: Boche, Holger, et al.
Publicado: (2024)
Overflow Prevention Enhances Long-Context Recurrent LLMs
por: Ben-Kish, Assaf, et al.
Publicado: (2025)
por: Ben-Kish, Assaf, et al.
Publicado: (2025)
Group Orthogonalization Regularization For Vision Models Adaptation and Robustness
por: Kurtz, Yoav, et al.
Publicado: (2023)
por: Kurtz, Yoav, et al.
Publicado: (2023)
Multiplicative Reweighting for Robust Neural Network Optimization
por: Bar, Noga, et al.
Publicado: (2021)
por: Bar, Noga, et al.
Publicado: (2021)
ParFam -- (Neural Guided) Symbolic Regression Based on Continuous Global Optimization
por: Scholl, Philipp, et al.
Publicado: (2023)
por: Scholl, Philipp, et al.
Publicado: (2023)
On the Relation Between Linear Diffusion and Power Iteration
por: Weitzner, Dana, et al.
Publicado: (2024)
por: Weitzner, Dana, et al.
Publicado: (2024)
Understanding Multimodal Failure in Action-Chunking Behavioral Cloning
por: Mazza, Lorenzo, et al.
Publicado: (2026)
por: Mazza, Lorenzo, et al.
Publicado: (2026)
Sparse-Aware Neural Networks for Nonlinear Functionals: Mitigating the Exponential Dependence on Dimension
por: Li, Jianfei, et al.
Publicado: (2026)
por: Li, Jianfei, et al.
Publicado: (2026)
Graph Hierarchical Recurrence for Long-Range Generalization
por: Carotti, Stefano, et al.
Publicado: (2026)
por: Carotti, Stefano, et al.
Publicado: (2026)
Revisiting Long-term Time Series Forecasting: An Investigation on Linear Mapping
por: Li, Zhe, et al.
Publicado: (2023)
por: Li, Zhe, et al.
Publicado: (2023)
Entropy-Preserving Reinforcement Learning
por: Petrenko, Aleksei, et al.
Publicado: (2026)
por: Petrenko, Aleksei, et al.
Publicado: (2026)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
por: Abuduweili, Abulikemu, et al.
Publicado: (2024)
When is a System Discoverable from Data? Discovery Requires Chaos
por: Shumaylov, Zakhar, et al.
Publicado: (2025)
por: Shumaylov, Zakhar, et al.
Publicado: (2025)
Can Continuous-Time Diffusion Models Generate and Solve Globally Constrained Discrete Problems? A Study on Sudoku
por: Drozdova, Mariia
Publicado: (2026)
por: Drozdova, Mariia
Publicado: (2026)
DeciMamba: Exploring the Length Extrapolation Potential of Mamba
por: Ben-Kish, Assaf, et al.
Publicado: (2024)
por: Ben-Kish, Assaf, et al.
Publicado: (2024)
CARES: Context-Aware Resolution Selector for VLMs
por: Kimhi, Moshe, et al.
Publicado: (2025)
por: Kimhi, Moshe, et al.
Publicado: (2025)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
por: Norman, Yotam, et al.
Publicado: (2025)
por: Norman, Yotam, et al.
Publicado: (2025)
Do multimodal models imagine electric sheep?
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
por: Ramakrishnan, Santhosh Kumar, et al.
Publicado: (2026)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
por: Yehudai, Asaf, et al.
Publicado: (2025)
por: Yehudai, Asaf, et al.
Publicado: (2025)
Linear Contextual Bandits with Hybrid Payoff: Revisited
por: Das, Nirjhar, et al.
Publicado: (2024)
por: Das, Nirjhar, et al.
Publicado: (2024)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
por: Ran-Milo, Yuval, et al.
Publicado: (2026)
Probabilistic neural operators for functional uncertainty quantification
por: Bülte, Christopher, et al.
Publicado: (2025)
por: Bülte, Christopher, et al.
Publicado: (2025)
Random Spiking Neural Networks are Stable and Spectrally Simple
por: Araya, Ernesto, et al.
Publicado: (2025)
por: Araya, Ernesto, et al.
Publicado: (2025)
Generalization Bounds for Message Passing Networks on Mixture of Graphons
por: Maskey, Sohir, et al.
Publicado: (2024)
por: Maskey, Sohir, et al.
Publicado: (2024)
Long-Range Graph Wavelet Networks
por: Guerranti, Filippo, et al.
Publicado: (2025)
por: Guerranti, Filippo, et al.
Publicado: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
por: Hu, Wenjie, et al.
Publicado: (2025)
por: Hu, Wenjie, et al.
Publicado: (2025)
JuStRank: Benchmarking LLM Judges for System Ranking
por: Gera, Ariel, et al.
Publicado: (2024)
por: Gera, Ariel, et al.
Publicado: (2024)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
por: Zhao, Yike, et al.
Publicado: (2026)
por: Zhao, Yike, et al.
Publicado: (2026)
Optimal Decay Spectra for Linear Recurrences
por: Cao, Yang
Publicado: (2026)
por: Cao, Yang
Publicado: (2026)
Provable Benefits of Complex Parameterizations for Structured State Space Models
por: Ran-Milo, Yuval, et al.
Publicado: (2024)
por: Ran-Milo, Yuval, et al.
Publicado: (2024)
Interpreting Affine Recurrence Learning in GPT-style Transformers
por: Bhargav, Samarth, et al.
Publicado: (2024)
por: Bhargav, Samarth, et al.
Publicado: (2024)
Block-Biased Mamba for Long-Range Sequence Processing
por: Yu, Annan, et al.
Publicado: (2025)
por: Yu, Annan, et al.
Publicado: (2025)
Provable Long-Range Benefits of Next-Token Prediction
por: Cao, Xinyuan, et al.
Publicado: (2025)
por: Cao, Xinyuan, et al.
Publicado: (2025)
On Measuring Long-Range Interactions in Graph Neural Networks
por: Bamberger, Jacob, et al.
Publicado: (2025)
por: Bamberger, Jacob, et al.
Publicado: (2025)
Ejemplares similares
-
How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences
por: Seleznova, Mariia
Publicado: (2026) -
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
por: Bar, Noga, et al.
Publicado: (2024) -
Pruning at Initialization -- A Sketching Perspective
por: Bar, Noga, et al.
Publicado: (2023) -
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
por: Razin, Noam, et al.
Publicado: (2024) -
ZOQO: Zero-Order Quantized Optimization
por: Bar, Noga, et al.
Publicado: (2025)