Saved in:
| Main Authors: | Irie, Kazuki, Gopalakrishnan, Anand, Schmidhuber, Jürgen |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2305.19044 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Metalearning Continual Learning Algorithms
by: Irie, Kazuki, et al.
Published: (2023)
by: Irie, Kazuki, et al.
Published: (2023)
Self-Organising Neural Discrete Representation Learning à la Kohonen
by: Irie, Kazuki, et al.
Published: (2023)
by: Irie, Kazuki, et al.
Published: (2023)
Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
by: Gopalakrishnan, Anand, et al.
Published: (2024)
by: Gopalakrishnan, Anand, et al.
Published: (2024)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
by: Csordás, Róbert, et al.
Published: (2023)
by: Csordás, Róbert, et al.
Published: (2023)
Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
by: Gopalakrishnan, Anand, et al.
Published: (2025)
by: Gopalakrishnan, Anand, et al.
Published: (2025)
MoEUT: Mixture-of-Experts Universal Transformers
by: Csordás, Róbert, et al.
Published: (2024)
by: Csordás, Róbert, et al.
Published: (2024)
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
by: Irie, Kazuki
Published: (2024)
by: Irie, Kazuki
Published: (2024)
Learning Useful Representations of Recurrent Neural Network Weight Matrices
by: Herrmann, Vincent, et al.
Published: (2024)
by: Herrmann, Vincent, et al.
Published: (2024)
Overcoming classic challenges for artificial neural networks by providing incentives and practice
by: Irie, Kazuki, et al.
Published: (2024)
by: Irie, Kazuki, et al.
Published: (2024)
Fast weight programming and linear transformers: from machine learning to neurobiology
by: Irie, Kazuki, et al.
Published: (2025)
by: Irie, Kazuki, et al.
Published: (2025)
Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
by: Irie, Kazuki, et al.
Published: (2025)
by: Irie, Kazuki, et al.
Published: (2025)
Learning to Forget: Continual Learning with Adaptive Weight Decay
by: Ramesh, Aditya A., et al.
Published: (2026)
by: Ramesh, Aditya A., et al.
Published: (2026)
Key-value memory in the brain
by: Gershman, Samuel J., et al.
Published: (2025)
by: Gershman, Samuel J., et al.
Published: (2025)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Interestingness as an Inductive Heuristic for Future Compression Progress
by: Herrmann, Vincent, et al.
Published: (2026)
by: Herrmann, Vincent, et al.
Published: (2026)
Real-Time Recurrent Reinforcement Learning
by: Lemmel, Julian, et al.
Published: (2023)
by: Lemmel, Julian, et al.
Published: (2023)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025)
by: Piękos, Piotr, et al.
Published: (2025)
Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View
by: Heo, Gyuryang, et al.
Published: (2026)
by: Heo, Gyuryang, et al.
Published: (2026)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
by: Tiberi, Lorenzo, et al.
Published: (2024)
by: Tiberi, Lorenzo, et al.
Published: (2024)
Who invented deep residual learning?
by: Schmidhuber, Juergen
Published: (2025)
by: Schmidhuber, Juergen
Published: (2025)
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
by: Kawakami, Tatsuki, et al.
Published: (2025)
by: Kawakami, Tatsuki, et al.
Published: (2025)
Sequence Compression Speeds Up Credit Assignment in Reinforcement Learning
by: Ramesh, Aditya A., et al.
Published: (2024)
by: Ramesh, Aditya A., et al.
Published: (2024)
Fairness Overfitting in Machine Learning: An Information-Theoretic Perspective
by: Laakom, Firas, et al.
Published: (2025)
by: Laakom, Firas, et al.
Published: (2025)
Measuring In-Context Computation Complexity via Hidden State Prediction
by: Herrmann, Vincent, et al.
Published: (2025)
by: Herrmann, Vincent, et al.
Published: (2025)
Streaming Reinforcement Learning under Partial Observability with Real-Time Recurrent Learning
by: Farr, Noah, et al.
Published: (2026)
by: Farr, Noah, et al.
Published: (2026)
Sequential-Parallel Duality in Prefix Scannable Models
by: Yau, Morris, et al.
Published: (2025)
by: Yau, Morris, et al.
Published: (2025)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Splats under Pressure: Exploring Performance-Energy Trade-offs in Real-Time 3D Gaussian Splatting under Constrained GPU Budgets
by: Tajwar, Muhammad Fahim, et al.
Published: (2026)
by: Tajwar, Muhammad Fahim, et al.
Published: (2026)
Multiple Token Divergence: Measuring and Steering In-Context Computation Density
by: Herrmann, Vincent, et al.
Published: (2025)
by: Herrmann, Vincent, et al.
Published: (2025)
Exploring the limits of Hierarchical World Models in Reinforcement Learning
by: Schiewer, Robin, et al.
Published: (2024)
by: Schiewer, Robin, et al.
Published: (2024)
Autonomous AI-based Cybersecurity Framework for Critical Infrastructure: Real-Time Threat Mitigation
by: Paulraj, Jenifer, et al.
Published: (2025)
by: Paulraj, Jenifer, et al.
Published: (2025)
FACTS: A Factored State-Space Framework For World Modelling
by: Nanbo, Li, et al.
Published: (2024)
by: Nanbo, Li, et al.
Published: (2024)
Highway Value Iteration Networks
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
Impact of Recurrent Neural Networks and Deep Learning Frameworks on Real-time Lightweight Time Series Anomaly Detection
by: Lee, Ming-Chang, et al.
Published: (2024)
by: Lee, Ming-Chang, et al.
Published: (2024)
Highway Reinforcement Learning
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
Curious Causality-Seeking Agents Learn Meta Causal World
by: Zhao, Zhiyu, et al.
Published: (2025)
by: Zhao, Zhiyu, et al.
Published: (2025)
Fast and scalable retrosynthetic planning with a transformer neural network and speculative beam search
by: Andronov, Mikhail, et al.
Published: (2025)
by: Andronov, Mikhail, et al.
Published: (2025)
Large Engagement Networks for Classifying Coordinated Campaigns and Organic Twitter Trends
by: Gopalakrishnan, Atul Anand, et al.
Published: (2025)
by: Gopalakrishnan, Atul Anand, et al.
Published: (2025)
Density-aware Walks for Coordinated Campaign Detection
by: Gopalakrishnan, Atul Anand, et al.
Published: (2025)
by: Gopalakrishnan, Atul Anand, et al.
Published: (2025)
Time-Warping Recurrent Neural Networks for Transfer Learning
by: Hirschi, Jonathon
Published: (2026)
by: Hirschi, Jonathon
Published: (2026)
Similar Items
-
Metalearning Continual Learning Algorithms
by: Irie, Kazuki, et al.
Published: (2023) -
Self-Organising Neural Discrete Representation Learning à la Kohonen
by: Irie, Kazuki, et al.
Published: (2023) -
Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
by: Gopalakrishnan, Anand, et al.
Published: (2024) -
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
by: Csordás, Róbert, et al.
Published: (2023) -
Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
by: Gopalakrishnan, Anand, et al.
Published: (2025)