Multi-agent cooperation through learning-aware policy gradients
Fuente:
arXiv
Guardado en:
| Autores principales: | Meulemans, Alexander, Kobayashi, Seijin, von Oswald, Johannes, Scherrer, Nino, Elmoznino, Eric, Richards, Blake, Lajoie, Guillaume, Arcas, Blaise Agüera y, Sacramento, João |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
por: Kobayashi, Seijin, et al.
Publicado: (2025)
por: Kobayashi, Seijin, et al.
Publicado: (2025)
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
por: Meulemans, Alexander, et al.
Publicado: (2025)
por: Meulemans, Alexander, et al.
Publicado: (2025)
Uncovering mesa-optimization algorithms in Transformers
por: von Oswald, Johannes, et al.
Publicado: (2023)
por: von Oswald, Johannes, et al.
Publicado: (2023)
Multi-agent cooperation through in-context co-player inference
por: Weis, Marissa A., et al.
Publicado: (2026)
por: Weis, Marissa A., et al.
Publicado: (2026)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
por: von Oswald, Johannes, et al.
Publicado: (2025)
por: von Oswald, Johannes, et al.
Publicado: (2025)
When can transformers compositionally generalize in-context?
por: Kobayashi, Seijin, et al.
Publicado: (2024)
por: Kobayashi, Seijin, et al.
Publicado: (2024)
Reasoning Models Generate Societies of Thought
por: Kim, Junsol, et al.
Publicado: (2026)
por: Kim, Junsol, et al.
Publicado: (2026)
The unreasonable effectiveness of pattern matching
por: Lupyan, Gary, et al.
Publicado: (2026)
por: Lupyan, Gary, et al.
Publicado: (2026)
Gated recurrent neural networks discover attention
por: Zucchet, Nicolas, et al.
Publicado: (2023)
por: Zucchet, Nicolas, et al.
Publicado: (2023)
Agentic AI and the next intelligence explosion
por: Evans, James, et al.
Publicado: (2026)
por: Evans, James, et al.
Publicado: (2026)
Learning Randomized Algorithms with Transformers
por: von Oswald, Johannes, et al.
Publicado: (2024)
por: von Oswald, Johannes, et al.
Publicado: (2024)
Weight decay induces low-rank attention layers
por: Kobayashi, Seijin, et al.
Publicado: (2024)
por: Kobayashi, Seijin, et al.
Publicado: (2024)
A Complexity-Based Theory of Compositionality
por: Elmoznino, Eric, et al.
Publicado: (2024)
por: Elmoznino, Eric, et al.
Publicado: (2024)
What Lives? A meta-analysis of diverse opinions on the definition of life
por: Bender, Reed, et al.
Publicado: (2025)
por: Bender, Reed, et al.
Publicado: (2025)
Discovering modular solutions that generalize compositionally
por: Schug, Simon, et al.
Publicado: (2023)
por: Schug, Simon, et al.
Publicado: (2023)
Attention as a Hypernetwork
por: Schug, Simon, et al.
Publicado: (2024)
por: Schug, Simon, et al.
Publicado: (2024)
Can LLMs get help from other LLMs without revealing private information?
por: Hartmann, Florian, et al.
Publicado: (2024)
por: Hartmann, Florian, et al.
Publicado: (2024)
Social Learning: Towards Collaborative Learning with Large Language Models
por: Mohtashami, Amirkeivan, et al.
Publicado: (2023)
por: Mohtashami, Amirkeivan, et al.
Publicado: (2023)
Discrete, compositional, and symbolic representations through attractor dynamics
por: Nam, Andrew, et al.
Publicado: (2023)
por: Nam, Andrew, et al.
Publicado: (2023)
Does learning the right latent variables necessarily improve in-context learning?
por: Mittal, Sarthak, et al.
Publicado: (2024)
por: Mittal, Sarthak, et al.
Publicado: (2024)
Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
por: Gagnon, Leo, et al.
Publicado: (2025)
por: Gagnon, Leo, et al.
Publicado: (2025)
State Soup: In-Context Skill Learning, Retrieval and Mixing
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
Amortizing intractable inference in large language models
por: Hu, Edward J., et al.
Publicado: (2023)
por: Hu, Edward J., et al.
Publicado: (2023)
Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
por: Kapl, Ferdinand, et al.
Publicado: (2025)
por: Kapl, Ferdinand, et al.
Publicado: (2025)
In-context learning and Occam's razor
por: Elmoznino, Eric, et al.
Publicado: (2024)
por: Elmoznino, Eric, et al.
Publicado: (2024)
Computational Life: How Well-formed, Self-replicating Programs Emerge from Simple Interaction
por: Arcas, Blaise Agüera y, et al.
Publicado: (2024)
por: Arcas, Blaise Agüera y, et al.
Publicado: (2024)
Synaptic Weight Distributions Depend on the Geometry of Plasticity
por: Pogodin, Roman, et al.
Publicado: (2023)
por: Pogodin, Roman, et al.
Publicado: (2023)
Sufficient conditions for offline reactivation in recurrent neural networks
por: Krishna, Nanda H., et al.
Publicado: (2025)
por: Krishna, Nanda H., et al.
Publicado: (2025)
A Compression Perspective on Simplicity Bias
por: Marty, Tom, et al.
Publicado: (2026)
por: Marty, Tom, et al.
Publicado: (2026)
Towards a future space-based, highly scalable AI infrastructure system design
por: Arcas, Blaise Agüera y, et al.
Publicado: (2025)
por: Arcas, Blaise Agüera y, et al.
Publicado: (2025)
The challenge of hidden gifts in multi-agent reinforcement learning
por: Malenfant, Dane, et al.
Publicado: (2025)
por: Malenfant, Dane, et al.
Publicado: (2025)
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
por: Firsching, Moritz, et al.
Publicado: (2026)
por: Firsching, Moritz, et al.
Publicado: (2026)
Can LLMs make trade-offs involving stipulated pain and pleasure states?
por: Keeling, Geoff, et al.
Publicado: (2024)
por: Keeling, Geoff, et al.
Publicado: (2024)
Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling
por: Rodchenko, Tanya, et al.
Publicado: (2025)
por: Rodchenko, Tanya, et al.
Publicado: (2025)
Engineering Sentience
por: Demin, Konstantin, et al.
Publicado: (2025)
por: Demin, Konstantin, et al.
Publicado: (2025)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
por: Li, Melody Zixuan, et al.
Publicado: (2025)
por: Li, Melody Zixuan, et al.
Publicado: (2025)
LLMs achieve adult human performance on higher-order theory of mind tasks
por: Street, Winnie, et al.
Publicado: (2024)
por: Street, Winnie, et al.
Publicado: (2024)
Lazy vs hasty: linearization in deep networks impacts learning schedule based on example difficulty
por: George, Thomas, et al.
Publicado: (2022)
por: George, Thomas, et al.
Publicado: (2022)
Beyond Distribution Sharpening: The Importance of Task Rewards
por: Mittal, Sarthak, et al.
Publicado: (2026)
por: Mittal, Sarthak, et al.
Publicado: (2026)
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
por: Williams, Ezekiel, et al.
Publicado: (2026)
por: Williams, Ezekiel, et al.
Publicado: (2026)
Ejemplares similares
-
Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
por: Kobayashi, Seijin, et al.
Publicado: (2025) -
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
por: Meulemans, Alexander, et al.
Publicado: (2025) -
Uncovering mesa-optimization algorithms in Transformers
por: von Oswald, Johannes, et al.
Publicado: (2023) -
Multi-agent cooperation through in-context co-player inference
por: Weis, Marissa A., et al.
Publicado: (2026) -
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
por: von Oswald, Johannes, et al.
Publicado: (2025)