MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
Fuente:
arXiv
Saved in:
| Main Authors: | von Oswald, Johannes, Scherrer, Nino, Kobayashi, Seijin, Versari, Luca, Yang, Songlin, Schlegel, Maximilian, Maile, Kaitlin, Schimpf, Yanick, Sieberling, Oliver, Meulemans, Alexander, Saurous, Rif A., Lajoie, Guillaume, Frenkel, Charlotte, Pascanu, Razvan, Arcas, Blaise Agüera y, Sacramento, João |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
by: Kobayashi, Seijin, et al.
Published: (2025)
by: Kobayashi, Seijin, et al.
Published: (2025)
Multi-agent cooperation through learning-aware policy gradients
by: Meulemans, Alexander, et al.
Published: (2024)
by: Meulemans, Alexander, et al.
Published: (2024)
Uncovering mesa-optimization algorithms in Transformers
by: von Oswald, Johannes, et al.
Published: (2023)
by: von Oswald, Johannes, et al.
Published: (2023)
Multi-agent cooperation through in-context co-player inference
by: Weis, Marissa A., et al.
Published: (2026)
by: Weis, Marissa A., et al.
Published: (2026)
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
by: Meulemans, Alexander, et al.
Published: (2025)
by: Meulemans, Alexander, et al.
Published: (2025)
When can transformers compositionally generalize in-context?
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
Reasoning Models Generate Societies of Thought
by: Kim, Junsol, et al.
Published: (2026)
by: Kim, Junsol, et al.
Published: (2026)
The unreasonable effectiveness of pattern matching
by: Lupyan, Gary, et al.
Published: (2026)
by: Lupyan, Gary, et al.
Published: (2026)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
by: Kapl, Ferdinand, et al.
Published: (2025)
by: Kapl, Ferdinand, et al.
Published: (2025)
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023)
by: Schug, Simon, et al.
Published: (2023)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
All Random Features Representations are Equivalent
by: Sernau, Luke, et al.
Published: (2024)
by: Sernau, Luke, et al.
Published: (2024)
Agentic AI and the next intelligence explosion
by: Evans, James, et al.
Published: (2026)
by: Evans, James, et al.
Published: (2026)
Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
by: Arcas, Blaise Agüera y, et al.
Published: (2024)
by: Arcas, Blaise Agüera y, et al.
Published: (2024)
What Lives? A meta-analysis of diverse opinions on the definition of life
by: Bender, Reed, et al.
Published: (2025)
by: Bender, Reed, et al.
Published: (2025)
From Growing to Looping: A Unified View of Iterative Computation in LLMs
by: Kapl, Ferdinand, et al.
Published: (2026)
by: Kapl, Ferdinand, et al.
Published: (2026)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Can LLMs get help from other LLMs without revealing private information?
by: Hartmann, Florian, et al.
Published: (2024)
by: Hartmann, Florian, et al.
Published: (2024)
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023)
by: Zucchet, Nicolas, et al.
Published: (2023)
Social Learning: Towards Collaborative Learning with Large Language Models
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
by: Mohtashami, Amirkeivan, et al.
Published: (2023)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Deep Grokking: Would Deep Neural Networks Generalize Better?
by: Fan, Simin, et al.
Published: (2024)
by: Fan, Simin, et al.
Published: (2024)
Scalable Spatiotemporal Prediction with Bayesian Neural Fields
by: Saad, Feras, et al.
Published: (2024)
by: Saad, Feras, et al.
Published: (2024)
Robust Inverse Graphics via Probabilistic Inference
by: Le, Tuan Anh, et al.
Published: (2024)
by: Le, Tuan Anh, et al.
Published: (2024)
Learning Randomized Algorithms with Transformers
by: von Oswald, Johannes, et al.
Published: (2024)
by: von Oswald, Johannes, et al.
Published: (2024)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
by: Li, Qinyu, et al.
Published: (2025)
by: Li, Qinyu, et al.
Published: (2025)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
What Can Grokking Teach Us About Learning Under Nonstationarity?
by: Lyle, Clare, et al.
Published: (2025)
by: Lyle, Clare, et al.
Published: (2025)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
Softmax is not Enough (for Sharp Size Generalisation)
by: Veličković, Petar, et al.
Published: (2024)
by: Veličković, Petar, et al.
Published: (2024)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
Evolution With Purpose: Hierarchy-Informed Optimization of Whole-Brain Models
by: Shahrzad, Hormoz, et al.
Published: (2026)
by: Shahrzad, Hormoz, et al.
Published: (2026)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling
by: Rodchenko, Tanya, et al.
Published: (2025)
by: Rodchenko, Tanya, et al.
Published: (2025)
anw-sh/ind_tb_reanalysis: ind_tb_reanalysis
by: Anwesh Maile
Published: (2025)
by: Anwesh Maile
Published: (2025)
Similar Items
-
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
by: Kobayashi, Seijin, et al.
Published: (2025) -
Multi-agent cooperation through learning-aware policy gradients
by: Meulemans, Alexander, et al.
Published: (2024) -
Uncovering mesa-optimization algorithms in Transformers
by: von Oswald, Johannes, et al.
Published: (2023) -
Multi-agent cooperation through in-context co-player inference
by: Weis, Marissa A., et al.
Published: (2026) -
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
by: Meulemans, Alexander, et al.
Published: (2025)