Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Irie, Kazuki, Yau, Morris, Gershman, Samuel J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fast weight programming and linear transformers: from machine learning to neurobiology
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
von: Irie, Kazuki, et al.
Veröffentlicht: (2025)
Key-value memory in the brain
von: Gershman, Samuel J., et al.
Veröffentlicht: (2025)
von: Gershman, Samuel J., et al.
Veröffentlicht: (2025)
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
von: Heo, Gyuryang, et al.
Veröffentlicht: (2026)
von: Heo, Gyuryang, et al.
Veröffentlicht: (2026)
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
von: Irie, Kazuki
Veröffentlicht: (2024)
von: Irie, Kazuki
Veröffentlicht: (2024)
Sequential-Parallel Duality in Prefix Scannable Models
von: Yau, Morris, et al.
Veröffentlicht: (2025)
von: Yau, Morris, et al.
Veröffentlicht: (2025)
Metalearning Continual Learning Algorithms
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
Exploring the Promise and Limits of Real-Time Recurrent Learning
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
von: Csordás, Róbert, et al.
Veröffentlicht: (2023)
Overcoming classic challenges for artificial neural networks by providing incentives and practice
von: Irie, Kazuki, et al.
Veröffentlicht: (2024)
von: Irie, Kazuki, et al.
Veröffentlicht: (2024)
Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
von: Vastola, John J., et al.
Veröffentlicht: (2025)
von: Vastola, John J., et al.
Veröffentlicht: (2025)
A Variational Manifold Embedding Framework for Nonlinear Dimensionality Reduction
von: Vastola, John J., et al.
Veröffentlicht: (2025)
von: Vastola, John J., et al.
Veröffentlicht: (2025)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
von: Tiberi, Lorenzo, et al.
Veröffentlicht: (2024)
von: Tiberi, Lorenzo, et al.
Veröffentlicht: (2024)
Artificial intelligence for science: The easy and hard problems
von: Battleday, Ruairidh M., et al.
Veröffentlicht: (2024)
von: Battleday, Ruairidh M., et al.
Veröffentlicht: (2024)
Self-Organising Neural Discrete Representation Learning à la Kohonen
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
von: Irie, Kazuki, et al.
Veröffentlicht: (2023)
A circuit for predicting hierarchical structure in-context in Large Language Models
von: Saanum, Tankred, et al.
Veröffentlicht: (2025)
von: Saanum, Tankred, et al.
Veröffentlicht: (2025)
General Intelligence Requires Reward-based Pretraining
von: Han, Seungwook, et al.
Veröffentlicht: (2025)
von: Han, Seungwook, et al.
Veröffentlicht: (2025)
Successor-Predecessor Intrinsic Exploration
von: Yu, Changmin, et al.
Veröffentlicht: (2023)
von: Yu, Changmin, et al.
Veröffentlicht: (2023)
MoEUT: Mixture-of-Experts Universal Transformers
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
von: Csordás, Róbert, et al.
Veröffentlicht: (2024)
Hybrid Memory Replay: Blending Real and Distilled Data for Class Incremental Learning
von: Kong, Jiangtao, et al.
Veröffentlicht: (2024)
von: Kong, Jiangtao, et al.
Veröffentlicht: (2024)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
von: Khasia, Vladimer
Veröffentlicht: (2026)
von: Khasia, Vladimer
Veröffentlicht: (2026)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
von: Kawakami, Tatsuki, et al.
Veröffentlicht: (2025)
Grokking as the Transition from Lazy to Rich Training Dynamics
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2023)
Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems
von: Ziemann, Ingvar, et al.
Veröffentlicht: (2022)
von: Ziemann, Ingvar, et al.
Veröffentlicht: (2022)
Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
von: Carvalho, Wilka, et al.
Veröffentlicht: (2025)
von: Carvalho, Wilka, et al.
Veröffentlicht: (2025)
Principled Data Augmentation for Learning to Solve Quadratic Programming Problems
von: Qian, Chendi, et al.
Veröffentlicht: (2025)
von: Qian, Chendi, et al.
Veröffentlicht: (2025)
Any-Time Regret-Guaranteed Algorithm for Control of Linear Quadratic Systems
von: Chekan, Jafar Abbaszadeh, et al.
Veröffentlicht: (2024)
von: Chekan, Jafar Abbaszadeh, et al.
Veröffentlicht: (2024)
Predictive representations: building blocks of intelligence
von: Carvalho, Wilka, et al.
Veröffentlicht: (2024)
von: Carvalho, Wilka, et al.
Veröffentlicht: (2024)
Do Mice Grok? Glimpses of Hidden Progress During Overtraining in Sensory Cortex
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Variational Linear Attention: Stable Associative Memory for Long-Context Transformers
von: Pandey, Vishal, et al.
Veröffentlicht: (2026)
von: Pandey, Vishal, et al.
Veröffentlicht: (2026)
The Confusing Instance Principle for Online Linear Quadratic Control
von: Radji, Waris, et al.
Veröffentlicht: (2025)
von: Radji, Waris, et al.
Veröffentlicht: (2025)
Improving the Linearized Laplace Approximation via Quadratic Approximations
von: Jiménez, Pedro, et al.
Veröffentlicht: (2026)
von: Jiménez, Pedro, et al.
Veröffentlicht: (2026)
Learning Linear Attention in Polynomial Time
von: Yau, Morris, et al.
Veröffentlicht: (2024)
von: Yau, Morris, et al.
Veröffentlicht: (2024)
On the Expressive Power of GNNs to Solve Linear SDPs
von: Qian, Chendi, et al.
Veröffentlicht: (2026)
von: Qian, Chendi, et al.
Veröffentlicht: (2026)
Generalizing Linear Autoencoder Recommenders with Decoupled Expected Quadratic Loss
von: Guo, Ruixin, et al.
Veröffentlicht: (2026)
von: Guo, Ruixin, et al.
Veröffentlicht: (2026)
Latent Linear Quadratic Regulator for Robotic Control Tasks
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
Scalar Federated Learning for Linear Quadratic Regulator
von: Rostami, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Rostami, Mohammadreza, et al.
Veröffentlicht: (2026)
Accelerated Optimization Landscape of Linear-Quadratic Regulator
von: Feng, Lechen, et al.
Veröffentlicht: (2023)
von: Feng, Lechen, et al.
Veröffentlicht: (2023)
Task diversity produces systematic transfer but inhibits continual reinforcement learning
von: Seth, Purab, et al.
Veröffentlicht: (2026)
von: Seth, Purab, et al.
Veröffentlicht: (2026)
Covariance-Aware Transformers for Quadratic Programming and Decision Making
von: Tire, Kutay, et al.
Veröffentlicht: (2026)
von: Tire, Kutay, et al.
Veröffentlicht: (2026)
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
von: Ding, Shihong, et al.
Veröffentlicht: (2025)
von: Ding, Shihong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fast weight programming and linear transformers: from machine learning to neurobiology
von: Irie, Kazuki, et al.
Veröffentlicht: (2025) -
Key-value memory in the brain
von: Gershman, Samuel J., et al.
Veröffentlicht: (2025) -
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
von: Heo, Gyuryang, et al.
Veröffentlicht: (2026) -
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
von: Irie, Kazuki
Veröffentlicht: (2024) -
Sequential-Parallel Duality in Prefix Scannable Models
von: Yau, Morris, et al.
Veröffentlicht: (2025)