Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
Fuente:
arXiv
Guardado en:
| Autores principales: | Heo, Gyuryang, Ngotiaoco, Timothy, Irie, Kazuki, Gershman, Samuel J., Sabatini, Bernardo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fast weight programming and linear transformers: from machine learning to neurobiology
por: Irie, Kazuki, et al.
Publicado: (2025)
por: Irie, Kazuki, et al.
Publicado: (2025)
Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
por: Irie, Kazuki, et al.
Publicado: (2025)
por: Irie, Kazuki, et al.
Publicado: (2025)
Key-value memory in the brain
por: Gershman, Samuel J., et al.
Publicado: (2025)
por: Gershman, Samuel J., et al.
Publicado: (2025)
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
por: Irie, Kazuki
Publicado: (2024)
por: Irie, Kazuki
Publicado: (2024)
Why Are Linear RNNs More Parallelizable?
por: Merrill, William, et al.
Publicado: (2026)
por: Merrill, William, et al.
Publicado: (2026)
Metalearning Continual Learning Algorithms
por: Irie, Kazuki, et al.
Publicado: (2023)
por: Irie, Kazuki, et al.
Publicado: (2023)
Exploring the Promise and Limits of Real-Time Recurrent Learning
por: Irie, Kazuki, et al.
Publicado: (2023)
por: Irie, Kazuki, et al.
Publicado: (2023)
A circuit for predicting hierarchical structure in-context in Large Language Models
por: Saanum, Tankred, et al.
Publicado: (2025)
por: Saanum, Tankred, et al.
Publicado: (2025)
A Variational Manifold Embedding Framework for Nonlinear Dimensionality Reduction
por: Vastola, John J., et al.
Publicado: (2025)
por: Vastola, John J., et al.
Publicado: (2025)
Bregman Conditional Random Fields: Sequence Labeling with Parallelizable Inference Algorithms
por: Corro, Caio, et al.
Publicado: (2025)
por: Corro, Caio, et al.
Publicado: (2025)
Overcoming classic challenges for artificial neural networks by providing incentives and practice
por: Irie, Kazuki, et al.
Publicado: (2024)
por: Irie, Kazuki, et al.
Publicado: (2024)
Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
por: Vastola, John J., et al.
Publicado: (2025)
por: Vastola, John J., et al.
Publicado: (2025)
Parallelizable memory recurrent units
por: De Geeter, Florent, et al.
Publicado: (2026)
por: De Geeter, Florent, et al.
Publicado: (2026)
Artificial intelligence for science: The easy and hard problems
por: Battleday, Ruairidh M., et al.
Publicado: (2024)
por: Battleday, Ruairidh M., et al.
Publicado: (2024)
Self-Organising Neural Discrete Representation Learning à la Kohonen
por: Irie, Kazuki, et al.
Publicado: (2023)
por: Irie, Kazuki, et al.
Publicado: (2023)
Successor-Predecessor Intrinsic Exploration
por: Yu, Changmin, et al.
Publicado: (2023)
por: Yu, Changmin, et al.
Publicado: (2023)
General Intelligence Requires Reward-based Pretraining
por: Han, Seungwook, et al.
Publicado: (2025)
por: Han, Seungwook, et al.
Publicado: (2025)
Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information
por: Safaai, Houman, et al.
Publicado: (2026)
por: Safaai, Houman, et al.
Publicado: (2026)
Sequential-Parallel Duality in Prefix Scannable Models
por: Yau, Morris, et al.
Publicado: (2025)
por: Yau, Morris, et al.
Publicado: (2025)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
por: Kawakami, Tatsuki, et al.
Publicado: (2025)
por: Kawakami, Tatsuki, et al.
Publicado: (2025)
$\texttt{immrax}$: A Parallelizable and Differentiable Toolbox for Interval Analysis and Mixed Monotone Reachability in JAX
por: Harapanahalli, Akash, et al.
Publicado: (2024)
por: Harapanahalli, Akash, et al.
Publicado: (2024)
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
por: Bell, Samuel J., et al.
Publicado: (2024)
por: Bell, Samuel J., et al.
Publicado: (2024)
Any-Order Flexible Length Masked Diffusion
por: Kim, Jaeyeon, et al.
Publicado: (2025)
por: Kim, Jaeyeon, et al.
Publicado: (2025)
Mini-Hes: A Parallelizable Second-order Latent Factor Analysis Model
por: Wang, Jialiang, et al.
Publicado: (2024)
por: Wang, Jialiang, et al.
Publicado: (2024)
Credit Assignment via Neural Manifold Noise Correlation
por: Kang, Byungwoo, et al.
Publicado: (2026)
por: Kang, Byungwoo, et al.
Publicado: (2026)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
por: Csordás, Róbert, et al.
Publicado: (2023)
por: Csordás, Róbert, et al.
Publicado: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
por: Kumar, Tanishq, et al.
Publicado: (2023)
por: Kumar, Tanishq, et al.
Publicado: (2023)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
por: Tiberi, Lorenzo, et al.
Publicado: (2024)
por: Tiberi, Lorenzo, et al.
Publicado: (2024)
ParaRNN: An Interpretable and Parallelizable Recurrent Neural Network for Time-Dependent Data
por: Cai, Yuxi, et al.
Publicado: (2026)
por: Cai, Yuxi, et al.
Publicado: (2026)
Preemptive Solving of Future Problems: Multitask Preplay in Humans and Machines
por: Carvalho, Wilka, et al.
Publicado: (2025)
por: Carvalho, Wilka, et al.
Publicado: (2025)
A Near-optimal, Scalable and Parallelizable Framework for Stochastic Bandits Robust to Adversarial Corruptions and Beyond
por: Hu, Zicheng, et al.
Publicado: (2025)
por: Hu, Zicheng, et al.
Publicado: (2025)
An Algebraic View of the Expressivity of Recurrent Language Models
por: Nowak, Franz, et al.
Publicado: (2026)
por: Nowak, Franz, et al.
Publicado: (2026)
Predictive representations: building blocks of intelligence
por: Carvalho, Wilka, et al.
Publicado: (2024)
por: Carvalho, Wilka, et al.
Publicado: (2024)
Do Mice Grok? Glimpses of Hidden Progress During Overtraining in Sensory Cortex
por: Kumar, Tanishq, et al.
Publicado: (2024)
por: Kumar, Tanishq, et al.
Publicado: (2024)
Unlocking the Power of Boltzmann Machines by Parallelizable Sampler and Efficient Temperature Estimation
por: Kubo, Kentaro, et al.
Publicado: (2025)
por: Kubo, Kentaro, et al.
Publicado: (2025)
Why Domain Generalization Fail? A View of Necessity and Sufficiency
por: Vuong, Long-Tung, et al.
Publicado: (2025)
por: Vuong, Long-Tung, et al.
Publicado: (2025)
Interpretable Machine Learning for Spatial Science: A Lie-Algebraic Kernel for Rotationally Anisotropic Gaussian Processes
por: Warrior, Kane, et al.
Publicado: (2026)
por: Warrior, Kane, et al.
Publicado: (2026)
Practical and Parallelizable Algorithms for Non-Monotone Submodular Maximization with Size Constraint
por: Chen, Yixin, et al.
Publicado: (2020)
por: Chen, Yixin, et al.
Publicado: (2020)
When Explanations Lie: Why Many Modified BP Attributions Fail
por: Sixt, Leon, et al.
Publicado: (2019)
por: Sixt, Leon, et al.
Publicado: (2019)
Transport of Algebraic Structure to Latent Embeddings
por: Pfrommer, Samuel, et al.
Publicado: (2024)
por: Pfrommer, Samuel, et al.
Publicado: (2024)
Ejemplares similares
-
Fast weight programming and linear transformers: from machine learning to neurobiology
por: Irie, Kazuki, et al.
Publicado: (2025) -
Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
por: Irie, Kazuki, et al.
Publicado: (2025) -
Key-value memory in the brain
por: Gershman, Samuel J., et al.
Publicado: (2025) -
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
por: Irie, Kazuki
Publicado: (2024) -
Why Are Linear RNNs More Parallelizable?
por: Merrill, William, et al.
Publicado: (2026)