Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Kobayashi, Seijin, Schimpf, Yanick, Schlegel, Maximilian, Steger, Angelika, Wolczyk, Maciej, von Oswald, Johannes, Scherrer, Nino, Maile, Kaitlin, Lajoie, Guillaume, Richards, Blake A., Saurous, Rif A., Manyika, James, Arcas, Blaise Agüera y, Meulemans, Alexander, Sacramento, João |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
por: von Oswald, Johannes, et al.
Publicado: (2025)
por: von Oswald, Johannes, et al.
Publicado: (2025)
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
por: Meulemans, Alexander, et al.
Publicado: (2025)
por: Meulemans, Alexander, et al.
Publicado: (2025)
Multi-agent cooperation through learning-aware policy gradients
por: Meulemans, Alexander, et al.
Publicado: (2024)
por: Meulemans, Alexander, et al.
Publicado: (2024)
Multi-agent cooperation through in-context co-player inference
por: Weis, Marissa A., et al.
Publicado: (2026)
por: Weis, Marissa A., et al.
Publicado: (2026)
Uncovering mesa-optimization algorithms in Transformers
por: von Oswald, Johannes, et al.
Publicado: (2023)
por: von Oswald, Johannes, et al.
Publicado: (2023)
Discovering modular solutions that generalize compositionally
por: Schug, Simon, et al.
Publicado: (2023)
por: Schug, Simon, et al.
Publicado: (2023)
Learning Randomized Algorithms with Transformers
por: von Oswald, Johannes, et al.
Publicado: (2024)
por: von Oswald, Johannes, et al.
Publicado: (2024)
Reasoning Models Generate Societies of Thought
por: Kim, Junsol, et al.
Publicado: (2026)
por: Kim, Junsol, et al.
Publicado: (2026)
Gated recurrent neural networks discover attention
por: Zucchet, Nicolas, et al.
Publicado: (2023)
por: Zucchet, Nicolas, et al.
Publicado: (2023)
The unreasonable effectiveness of pattern matching
por: Lupyan, Gary, et al.
Publicado: (2026)
por: Lupyan, Gary, et al.
Publicado: (2026)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
por: Kapl, Ferdinand, et al.
Publicado: (2025)
por: Kapl, Ferdinand, et al.
Publicado: (2025)
All Random Features Representations are Equivalent
por: Sernau, Luke, et al.
Publicado: (2024)
por: Sernau, Luke, et al.
Publicado: (2024)
Agentic AI and the next intelligence explosion
por: Evans, James, et al.
Publicado: (2026)
por: Evans, James, et al.
Publicado: (2026)
When can transformers compositionally generalize in-context?
por: Kobayashi, Seijin, et al.
Publicado: (2024)
por: Kobayashi, Seijin, et al.
Publicado: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
What Lives? A meta-analysis of diverse opinions on the definition of life
por: Bender, Reed, et al.
Publicado: (2025)
por: Bender, Reed, et al.
Publicado: (2025)
Towards a future space-based, highly scalable AI infrastructure system design
por: Arcas, Blaise Agüera y, et al.
Publicado: (2025)
por: Arcas, Blaise Agüera y, et al.
Publicado: (2025)
From Growing to Looping: A Unified View of Iterative Computation in LLMs
por: Kapl, Ferdinand, et al.
Publicado: (2026)
por: Kapl, Ferdinand, et al.
Publicado: (2026)
Weight decay induces low-rank attention layers
por: Kobayashi, Seijin, et al.
Publicado: (2024)
por: Kobayashi, Seijin, et al.
Publicado: (2024)
Can LLMs get help from other LLMs without revealing private information?
por: Hartmann, Florian, et al.
Publicado: (2024)
por: Hartmann, Florian, et al.
Publicado: (2024)
Social Learning: Towards Collaborative Learning with Large Language Models
por: Mohtashami, Amirkeivan, et al.
Publicado: (2023)
por: Mohtashami, Amirkeivan, et al.
Publicado: (2023)
Scalable Spatiotemporal Prediction with Bayesian Neural Fields
por: Saad, Feras, et al.
Publicado: (2024)
por: Saad, Feras, et al.
Publicado: (2024)
Robust Inverse Graphics via Probabilistic Inference
por: Le, Tuan Anh, et al.
Publicado: (2024)
por: Le, Tuan Anh, et al.
Publicado: (2024)
Attention as a Hypernetwork
por: Schug, Simon, et al.
Publicado: (2024)
por: Schug, Simon, et al.
Publicado: (2024)
Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
por: Arcas, Blaise Agüera y, et al.
Publicado: (2024)
por: Arcas, Blaise Agüera y, et al.
Publicado: (2024)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
por: Emmons, Scott, et al.
Publicado: (2025)
por: Emmons, Scott, et al.
Publicado: (2025)
Periodic autoregressive moving average models for the analysis of streamflow [abstract]
por: Slack, J.R.
Publicado: (1988)
por: Slack, J.R.
Publicado: (1988)
Evolution With Purpose: Hierarchy-Informed Optimization of Whole-Brain Models
por: Shahrzad, Hormoz, et al.
Publicado: (2026)
por: Shahrzad, Hormoz, et al.
Publicado: (2026)
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling
por: Rodchenko, Tanya, et al.
Publicado: (2025)
por: Rodchenko, Tanya, et al.
Publicado: (2025)
anw-sh/ind_tb_reanalysis: ind_tb_reanalysis
por: Anwesh Maile
Publicado: (2025)
por: Anwesh Maile
Publicado: (2025)
Managing Change
por: Maile, Michael
Publicado: (2024)
por: Maile, Michael
Publicado: (2024)
La casa de convalecencia / Maile Chapman ; traducción de Magdalena Palmer
por: Chapman, Maile
Publicado: (2011)
por: Chapman, Maile
Publicado: (2011)
Factores de riesgo relevantes asociados a las malformaciones congénitas en la provincia de Cienfuegos, 2008-2013
por: Mailé Santos Solís
Publicado: (2016)
por: Mailé Santos Solís
Publicado: (2016)
Niveles para la capacitación en una organización
por: Mailé Salgado-Cruz
Publicado: (2017)
por: Mailé Salgado-Cruz
Publicado: (2017)
Can LLMs make trade-offs involving stipulated pain and pleasure states?
por: Keeling, Geoff, et al.
Publicado: (2024)
por: Keeling, Geoff, et al.
Publicado: (2024)
Survival and growth of local and transplanted blue mussels (Mytilus trossulus, Lamark). / Jenia F. Yanick
por: Yanick, Jenia F
Publicado: (2003)
por: Yanick, Jenia F
Publicado: (2003)
Charge‐Mediated Interactions Affect Enzymatic Reactions in Peptide Condensates
por: Rif Harris, et al.
Publicado: (2024)
por: Rif Harris, et al.
Publicado: (2024)
A novel reference prior for Gaussian hierarchical models with intrinsic conditional autoregressive random effects
por: Ferreira, Marco A. R.
Publicado: (2026)
por: Ferreira, Marco A. R.
Publicado: (2026)
Theory of Existing: The Metabolism of Reality Through Harmonic Processing
por: Steger, Cynthia
Publicado: (2026)
por: Steger, Cynthia
Publicado: (2026)
A Framework for AI Safety Through Accurate System Description
por: Steger, Cynthia
Publicado: (2025)
por: Steger, Cynthia
Publicado: (2025)
Ejemplares similares
-
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
por: von Oswald, Johannes, et al.
Publicado: (2025) -
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
por: Meulemans, Alexander, et al.
Publicado: (2025) -
Multi-agent cooperation through learning-aware policy gradients
por: Meulemans, Alexander, et al.
Publicado: (2024) -
Multi-agent cooperation through in-context co-player inference
por: Weis, Marissa A., et al.
Publicado: (2026) -
Uncovering mesa-optimization algorithms in Transformers
por: von Oswald, Johannes, et al.
Publicado: (2023)