Dissecting Linear Recurrent Models: How Different Gating Strategies Drive Selectivity and Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Bouhadjar, Younes, Fabre, Maxime, Schmidt, Felix, Neftci, Emre |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiLIF: Structured State Space Model Dynamics and Parametrization for Spiking Neural Networks
by: Fabre, Maxime, et al.
Published: (2025)
by: Fabre, Maxime, et al.
Published: (2025)
SymSeqBench: a unified framework for the generation and analysis of rule-based symbolic sequences and datasets
by: Zajzon, Barna, et al.
Published: (2025)
by: Zajzon, Barna, et al.
Published: (2025)
Sparse Axonal and Dendritic Delays Enable Competitive SNNs for Keyword Classification
by: Bouhadjar, Younes, et al.
Published: (2026)
by: Bouhadjar, Younes, et al.
Published: (2026)
QS4D: Quantization-aware training for efficient hardware deployment of structured state-space sequential models
by: Siegel, Sebastian, et al.
Published: (2025)
by: Siegel, Sebastian, et al.
Published: (2025)
Zero-Shot Temporal Resolution Domain Adaptation for Spiking Neural Networks
by: Karilanova, Sanja, et al.
Published: (2024)
by: Karilanova, Sanja, et al.
Published: (2024)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
by: De, Soham, et al.
Published: (2024)
by: De, Soham, et al.
Published: (2024)
Unsupervised Learning of Spatio-Temporal Patterns in Spiking Neuronal Networks
by: Feiler, Florian, et al.
Published: (2024)
by: Feiler, Florian, et al.
Published: (2024)
Liger: Linearizing Large Language Models to Gated Recurrent Structures
by: Lan, Disen, et al.
Published: (2025)
by: Lan, Disen, et al.
Published: (2025)
Optimal Gradient Checkpointing for Sparse and Recurrent Architectures using Off-Chip Memory
by: Bencheikh, Wadjih, et al.
Published: (2024)
by: Bencheikh, Wadjih, et al.
Published: (2024)
QS4D: Quantization‐Aware Training for Efficient Hardware Deployment of Structured State‐Space Sequential Models
by: Sebastian Siegel, et al.
Published: (2026)
by: Sebastian Siegel, et al.
Published: (2026)
Dissecting Language Models: Machine Unlearning via Selective Pruning
by: Pochinkov, Nicholas, et al.
Published: (2024)
by: Pochinkov, Nicholas, et al.
Published: (2024)
GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling
by: Katsch, Tobias
Published: (2023)
by: Katsch, Tobias
Published: (2023)
Masked Gated Linear Unit
by: Tajima, Yukito, et al.
Published: (2025)
by: Tajima, Yukito, et al.
Published: (2025)
Text Sentiment Analysis and Classification Based on Bidirectional Gated Recurrent Units (GRUs) Model
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
Gated Linear Attention Transformers with Hardware-Efficient Training
by: Yang, Songlin, et al.
Published: (2023)
by: Yang, Songlin, et al.
Published: (2023)
SNNAX -- Spiking Neural Networks in JAX
by: Lohoff, Jamie, et al.
Published: (2024)
by: Lohoff, Jamie, et al.
Published: (2024)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025)
by: Zheng, Chen, et al.
Published: (2025)
Revisiting Bi-Linear State Transitions in Recurrent Neural Networks
by: Ebrahimi, M. Reza, et al.
Published: (2025)
by: Ebrahimi, M. Reza, et al.
Published: (2025)
Advancing Regular Language Reasoning in Linear Recurrent Neural Networks
by: Fan, Ting-Han, et al.
Published: (2023)
by: Fan, Ting-Han, et al.
Published: (2023)
Dissecting Fine-Tuning Unlearning in Large Language Models
by: Hong, Yihuai, et al.
Published: (2024)
by: Hong, Yihuai, et al.
Published: (2024)
Enhancing Bangla Fake News Detection Using Bidirectional Gated Recurrent Units and Deep Learning Techniques
by: Roy, Utsha, et al.
Published: (2024)
by: Roy, Utsha, et al.
Published: (2024)
Optimal Decay Spectra for Linear Recurrences
by: Cao, Yang
Published: (2026)
by: Cao, Yang
Published: (2026)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
by: Schwethelm, Kristian, et al.
Published: (2026)
by: Schwethelm, Kristian, et al.
Published: (2026)
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
by: Mu, Yongyu, et al.
Published: (2025)
by: Mu, Yongyu, et al.
Published: (2025)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
by: Huang, Zeyi, et al.
Published: (2026)
by: Huang, Zeyi, et al.
Published: (2026)
A Truly Sparse and General Implementation of Gradient-Based Synaptic Plasticity
by: Lohoff, Jamie, et al.
Published: (2025)
by: Lohoff, Jamie, et al.
Published: (2025)
Investigating the Impact of Data Selection Strategies on Language Model Performance
by: Gu, Jiayao, et al.
Published: (2025)
by: Gu, Jiayao, et al.
Published: (2025)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
by: Zhou, Yujun, et al.
Published: (2025)
by: Zhou, Yujun, et al.
Published: (2025)
Optimizing Automatic Differentiation with Deep Reinforcement Learning
by: Lohoff, Jamie, et al.
Published: (2024)
by: Lohoff, Jamie, et al.
Published: (2024)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
Selective Generation for Controllable Language Models
by: Lee, Minjae, et al.
Published: (2023)
by: Lee, Minjae, et al.
Published: (2023)
Structured Recurrent Mixers for Massively Parallelized Sequence Generation
by: Badger, Benjamin L.
Published: (2026)
by: Badger, Benjamin L.
Published: (2026)
On the Representational Capacity of Recurrent Neural Language Models
by: Nowak, Franz, et al.
Published: (2023)
by: Nowak, Franz, et al.
Published: (2023)
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
by: Dong, Peijie, et al.
Published: (2026)
by: Dong, Peijie, et al.
Published: (2026)
Who is In Charge? Dissecting Role Conflicts in Instruction Following
by: Zeng, Siqi
Published: (2025)
by: Zeng, Siqi
Published: (2025)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Label-Efficient Model Selection for Text Generation
by: Ashury-Tahan, Shir, et al.
Published: (2024)
by: Ashury-Tahan, Shir, et al.
Published: (2024)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
Dissecting embedding method: learning higher-order structures from data
by: Tupikina, Liubov, et al.
Published: (2024)
by: Tupikina, Liubov, et al.
Published: (2024)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
by: Cheng, Yunfei, et al.
Published: (2024)
by: Cheng, Yunfei, et al.
Published: (2024)
Similar Items
-
SiLIF: Structured State Space Model Dynamics and Parametrization for Spiking Neural Networks
by: Fabre, Maxime, et al.
Published: (2025) -
SymSeqBench: a unified framework for the generation and analysis of rule-based symbolic sequences and datasets
by: Zajzon, Barna, et al.
Published: (2025) -
Sparse Axonal and Dendritic Delays Enable Competitive SNNs for Keyword Classification
by: Bouhadjar, Younes, et al.
Published: (2026) -
QS4D: Quantization-aware training for efficient hardware deployment of structured state-space sequential models
by: Siegel, Sebastian, et al.
Published: (2025) -
Zero-Shot Temporal Resolution Domain Adaptation for Spiking Neural Networks
by: Karilanova, Sanja, et al.
Published: (2024)