Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Aaditya K., Moskovitz, Ted, Dragutinovic, Sara, Hill, Felix, Chan, Stephanie C. Y., Saxe, Andrew M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
von: Dragutinović, Sara, et al.
Veröffentlicht: (2025)
von: Dragutinović, Sara, et al.
Veröffentlicht: (2025)
Training Dynamics of In-Context Learning in Linear Attention
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
HARP: A challenging human-annotated math reasoning benchmark
von: Yue, Albert S., et al.
Veröffentlicht: (2024)
von: Yue, Albert S., et al.
Veröffentlicht: (2024)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
von: Lee, Jin Hwa, et al.
Veröffentlicht: (2025)
von: Lee, Jin Hwa, et al.
Veröffentlicht: (2025)
To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
von: Dragutinović, Sara, et al.
Veröffentlicht: (2026)
von: Dragutinović, Sara, et al.
Veröffentlicht: (2026)
The broader spectrum of in-context learning
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
von: Zhang, Yedi, et al.
Veröffentlicht: (2025)
Machine Learning-Augmented Optimization of Large Bilevel and Two-stage Stochastic Programs: Application to Cycling Network Design
von: Chan, Timothy C. Y., et al.
Veröffentlicht: (2022)
von: Chan, Timothy C. Y., et al.
Veröffentlicht: (2022)
Meta-Learning Strategies through Value Maximization in Neural Networks
von: Carrasco-Davis, Rodrigo, et al.
Veröffentlicht: (2023)
von: Carrasco-Davis, Rodrigo, et al.
Veröffentlicht: (2023)
When Representations Align: Universality in Representation Learning Dynamics
von: van Rossem, Loek, et al.
Veröffentlicht: (2024)
von: van Rossem, Loek, et al.
Veröffentlicht: (2024)
Make Haste Slowly: A Theory of Emergent Structured Mixed Selectivity in Feature Learning ReLU Networks
von: Jarvis, Devon, et al.
Veröffentlicht: (2025)
von: Jarvis, Devon, et al.
Veröffentlicht: (2025)
Nonlinear dynamics of localization in neural receptive fields
von: Lufkin, Leon, et al.
Veröffentlicht: (2025)
von: Lufkin, Leon, et al.
Veröffentlicht: (2025)
Algorithm Development in Neural Networks: Insights from the Streaming Parity Task
von: van Rossem, Loek, et al.
Veröffentlicht: (2025)
von: van Rossem, Loek, et al.
Veröffentlicht: (2025)
Bayes' Power for Explaining In-Context Learning Generalizations
von: Müller, Samuel, et al.
Veröffentlicht: (2024)
von: Müller, Samuel, et al.
Veröffentlicht: (2024)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
Learned feature representations are biased by complexity, learning order, position, and more
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
Understanding Unimodal Bias in Multimodal Deep Linear Networks
von: Zhang, Yedi, et al.
Veröffentlicht: (2023)
von: Zhang, Yedi, et al.
Veröffentlicht: (2023)
When Are Bias-Free ReLU Networks Effectively Linear Networks?
von: Zhang, Yedi, et al.
Veröffentlicht: (2024)
von: Zhang, Yedi, et al.
Veröffentlicht: (2024)
CoopetitiveV: Leveraging LLM-powered Coopetitive Multi-Agent Prompting for High-quality Verilog Generation
von: Mi, Zhendong, et al.
Veröffentlicht: (2024)
von: Mi, Zhendong, et al.
Veröffentlicht: (2024)
Learning by Self-Explaining
von: Stammer, Wolfgang, et al.
Veröffentlicht: (2023)
von: Stammer, Wolfgang, et al.
Veröffentlicht: (2023)
Towards Provable Emergence of In-Context Reinforcement Learning
von: Wang, Jiuqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiuqi, et al.
Veröffentlicht: (2025)
Logarithmic Neyman Regret for Adaptive Estimation of the Average Treatment Effect
von: Neopane, Ojash, et al.
Veröffentlicht: (2024)
von: Neopane, Ojash, et al.
Veröffentlicht: (2024)
In-Context Learning Strategies Emerge Rationally
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
von: Wurgaft, Daniel, et al.
Veröffentlicht: (2025)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2025)
von: Sakamoto, Keitaro, et al.
Veröffentlicht: (2025)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
von: Xie, Zixuan, et al.
Veröffentlicht: (2026)
Emergence of In-Context Reinforcement Learning from Noise Distillation
von: Zisman, Ilya, et al.
Veröffentlicht: (2023)
von: Zisman, Ilya, et al.
Veröffentlicht: (2023)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
von: Yang, Liu, et al.
Veröffentlicht: (2025)
von: Yang, Liu, et al.
Veröffentlicht: (2025)
Language models show human-like content effects on reasoning tasks
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
Early learning of the optimal constant solution in neural networks and humans
von: Rubruck, Jirko, et al.
Veröffentlicht: (2024)
von: Rubruck, Jirko, et al.
Veröffentlicht: (2024)
Context and Diversity Matter: The Emergence of In-Context Learning in World Models
von: Wang, Fan, et al.
Veröffentlicht: (2025)
von: Wang, Fan, et al.
Veröffentlicht: (2025)
Optimistic Algorithms for Adaptive Estimation of the Average Treatment Effect
von: Neopane, Ojash, et al.
Veröffentlicht: (2025)
von: Neopane, Ojash, et al.
Veröffentlicht: (2025)
On The Specialization of Neural Modules
von: Jarvis, Devon, et al.
Veröffentlicht: (2024)
von: Jarvis, Devon, et al.
Veröffentlicht: (2024)
The emergence of sparse attention: impact of data distribution and benefits of repetition
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
von: Zucchet, Nicolas, et al.
Veröffentlicht: (2025)
Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
von: Dong, Yuxin, et al.
Veröffentlicht: (2025)
von: Dong, Yuxin, et al.
Veröffentlicht: (2025)
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
von: Dominé, Clémentine C. J., et al.
Veröffentlicht: (2024)
von: Dominé, Clémentine C. J., et al.
Veröffentlicht: (2024)
Representation biases: will we achieve complete understanding by analyzing representations?
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2025)
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2025)
Revisiting the Role of Relearning in Semantic Dementia
von: Jarvis, Devon, et al.
Veröffentlicht: (2025)
von: Jarvis, Devon, et al.
Veröffentlicht: (2025)
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
von: Njaradi, Valentina, et al.
Veröffentlicht: (2026)
von: Njaradi, Valentina, et al.
Veröffentlicht: (2026)
Transformers Don't In-Context Learn Least Squares Regression
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024) -
Softmax $\geq$ Linear: Transformers may learn to classify in-context by kernel gradient descent
von: Dragutinović, Sara, et al.
Veröffentlicht: (2025) -
Training Dynamics of In-Context Learning in Linear Attention
von: Zhang, Yedi, et al.
Veröffentlicht: (2025) -
HARP: A challenging human-annotated math reasoning benchmark
von: Yue, Albert S., et al.
Veröffentlicht: (2024) -
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
von: Lee, Jin Hwa, et al.
Veröffentlicht: (2025)