When can transformers compositionally generalize in-context?
Fuente:
arXiv
Saved in:
| Main Authors: | Kobayashi, Seijin, Schug, Simon, Akram, Yassir, Redhardt, Florian, von Oswald, Johannes, Pascanu, Razvan, Lajoie, Guillaume, Sacramento, João |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling can lead to compositional generalization
by: Redhardt, Florian, et al.
Published: (2025)
by: Redhardt, Florian, et al.
Published: (2025)
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023)
by: Schug, Simon, et al.
Published: (2023)
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023)
by: Zucchet, Nicolas, et al.
Published: (2023)
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024)
by: Schug, Simon, et al.
Published: (2024)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)
by: Orvieto, Antonio, et al.
Published: (2023)
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
by: Williams, Ezekiel, et al.
Published: (2026)
by: Williams, Ezekiel, et al.
Published: (2026)
Teaching signal synchronization in deep neural networks with prospective neurons
by: Zucchet, Nicoas, et al.
Published: (2025)
by: Zucchet, Nicoas, et al.
Published: (2025)
Synaptic Weight Distributions Depend on the Geometry of Plasticity
by: Pogodin, Roman, et al.
Published: (2023)
by: Pogodin, Roman, et al.
Published: (2023)
Sufficient conditions for offline reactivation in recurrent neural networks
by: Krishna, Nanda H., et al.
Published: (2025)
by: Krishna, Nanda H., et al.
Published: (2025)
State-space models can learn in-context by gradient descent
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
by: Sushma, Neeraj Mohan, et al.
Published: (2024)
How Sequential Algorithm Portfolios can benefit Black Box Optimization
by: Dinu, Catalin-Viorel, et al.
Published: (2026)
by: Dinu, Catalin-Viorel, et al.
Published: (2026)
When Large Language Model Meets Optimization
by: Huang, Sen, et al.
Published: (2024)
by: Huang, Sen, et al.
Published: (2024)
Short-reach Optical Communications: A Real-world Task for Neuromorphic Hardware
by: Arnold, Elias, et al.
Published: (2024)
by: Arnold, Elias, et al.
Published: (2024)
SQUAT: Stateful Quantization-Aware Training in Recurrent Spiking Neural Networks
by: Venkatesh, Sreyes, et al.
Published: (2024)
by: Venkatesh, Sreyes, et al.
Published: (2024)
How connectivity structure shapes rich and lazy learning in neural circuits
by: Liu, Yuhan Helena, et al.
Published: (2023)
by: Liu, Yuhan Helena, et al.
Published: (2023)
Evolutionary Algorithms Are Significantly More Robust to Noise When They Ignore It
by: Antipov, Denis, et al.
Published: (2024)
by: Antipov, Denis, et al.
Published: (2024)
PRIMETIME : Limits of LLMs in Temporal Primitives
by: Gaere, Edward, et al.
Published: (2025)
by: Gaere, Edward, et al.
Published: (2025)
GEEvo: Game Economy Generation and Balancing with Evolutionary Algorithms
by: Rupp, Florian, et al.
Published: (2024)
by: Rupp, Florian, et al.
Published: (2024)
Impact of spatial transformations on landscape features of CEC2022 basic benchmark problems
by: Yin, Haoran, et al.
Published: (2024)
by: Yin, Haoran, et al.
Published: (2024)
A Non-Dominated Sorting Evolutionary Algorithm Updating When Required
by: Farias, Lucas R. C., et al.
Published: (2025)
by: Farias, Lucas R. C., et al.
Published: (2025)
When to Truncate the Archive? On the Effect of the Truncation Frequency in Multi-Objective Optimisation
by: Cui, Zhiji, et al.
Published: (2025)
by: Cui, Zhiji, et al.
Published: (2025)
Energy-Efficient Implementation of Spiking Recurrent Cells on FPGA
by: Harmeling, Pascal, et al.
Published: (2026)
by: Harmeling, Pascal, et al.
Published: (2026)
Weight decay induces low-rank attention layers
by: Kobayashi, Seijin, et al.
Published: (2024)
by: Kobayashi, Seijin, et al.
Published: (2024)
Scaling Behaviors of Evolutionary Algorithms on GPUs: When Does Parallelism Pay Off?
by: Yu, Xinmeng, et al.
Published: (2026)
by: Yu, Xinmeng, et al.
Published: (2026)
Self-Adjusting Evolutionary Algorithms Are Slow on Multimodal Landscapes
by: Lengler, Johannes, et al.
Published: (2024)
by: Lengler, Johannes, et al.
Published: (2024)
Diversity-Preserving Exploitation of Crossover
by: Lengler, Johannes, et al.
Published: (2025)
by: Lengler, Johannes, et al.
Published: (2025)
Towards Evolutionary Optimization Using the Ising Model
by: Klüttermann, Simon
Published: (2025)
by: Klüttermann, Simon
Published: (2025)
When Spiking neural networks meet temporal attention image decoding and adaptive spiking neuron
by: Qiu, Xuerui, et al.
Published: (2024)
by: Qiu, Xuerui, et al.
Published: (2024)
Amortized Inference of Neuron Parameters on Analog Neuromorphic Hardware
by: Kaiser, Jakob, et al.
Published: (2026)
by: Kaiser, Jakob, et al.
Published: (2026)
Integrating programmable plasticity in experiment descriptions for analog neuromorphic hardware
by: Spilger, Philipp, et al.
Published: (2024)
by: Spilger, Philipp, et al.
Published: (2024)
Expressivity of Neural Networks with Random Weights and Learned Biases
by: Williams, Ezekiel, et al.
Published: (2024)
by: Williams, Ezekiel, et al.
Published: (2024)
Tight Runtime Bounds for Static Unary Unbiased Evolutionary Algorithms on Linear Functions
by: Doerr, Carola, et al.
Published: (2023)
by: Doerr, Carola, et al.
Published: (2023)
Near-Tight Runtime Guarantees for Many-Objective Evolutionary Algorithms
by: Wietheger, Simon, et al.
Published: (2024)
by: Wietheger, Simon, et al.
Published: (2024)
Reproduction of AdEx dynamics on neuromorphic hardware through data embedding and simulation-based inference
by: Huhle, Jakob, et al.
Published: (2024)
by: Huhle, Jakob, et al.
Published: (2024)
Simultaneous Model-Based Evolution of Constants and Expression Structure in GP-GOMEA for Symbolic Regression
by: Koch, Johannes, et al.
Published: (2026)
by: Koch, Johannes, et al.
Published: (2026)
Introns and Templates Matter: Rethinking Linkage in GP-GOMEA
by: Koch, Johannes, et al.
Published: (2026)
by: Koch, Johannes, et al.
Published: (2026)
A Robust, Open-Source Framework for Spiking Neural Networks on Low-End FPGAs
by: Fan, Andrew, et al.
Published: (2025)
by: Fan, Andrew, et al.
Published: (2025)
Real-time processing of analog signals on accelerated neuromorphic hardware
by: Stradmann, Yannik, et al.
Published: (2026)
by: Stradmann, Yannik, et al.
Published: (2026)
TONUS: Neuromorphic human pose estimation for artistic sound co-creation
by: Lecomte, Jules, et al.
Published: (2025)
by: Lecomte, Jules, et al.
Published: (2025)
Speeding Up the NSGA-II via Dynamic Population Sizes
by: Doerr, Benjamin, et al.
Published: (2025)
by: Doerr, Benjamin, et al.
Published: (2025)
Similar Items
-
Scaling can lead to compositional generalization
by: Redhardt, Florian, et al.
Published: (2025) -
Discovering modular solutions that generalize compositionally
by: Schug, Simon, et al.
Published: (2023) -
Gated recurrent neural networks discover attention
by: Zucchet, Nicolas, et al.
Published: (2023) -
Attention as a Hypernetwork
by: Schug, Simon, et al.
Published: (2024) -
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)